chartcoach
Reference

Catalog data

Inspect guideline entries, catalog locations, and the files in a published release.

Each row in entries.parquet is one guideline entry. This Parquet file stores tabular data for Python, JavaScript, and database readers. Entries contain labels, ordered sections, and BibTeX source records for citations.

Guideline entry record

This excerpt uses exact values from the quickstart's catalog. It shows the first section and complete reference:

{
  "id": "directly-label-series-instead-of-using-a-color-key",
  "title": "Directly label colored series instead of relying on a color key",
  "description": "For color-coded series identification, prefer direct labels on chart marks to improve readability and mitigate color-key decoding for readers with color-vision deficiency.",
  "labels": [
    "purpose:refine",
    "basis:heuristic",
    "quality:accessibility",
    "lever:text-annotation",
    "component:label:use",
    "component:legend:avoid",
    "needs:color-vision-deficiency",
    "polish:annotation"
  ],
  "sections": [
    {
      "role": "advice",
      "title": "Move labels onto the marks",
      "content": "Label colored series directly on the chart instead of making readers decode them through a color key. For example, place series names next to lines or areas, or label pie slices directly so readers do not have to match colors back and forth with a legend."
    }
  ],
  "references": [
    "@misc{muth_colorblindness_2020,\n author = {Lisa Charlotte Muth},\n howpublished = {\\url{https://www.datawrapper.de/blog/colorblindness-part2}},\n note = {Datawrapper Blog. Accessed: 2025-11-17},\n title = {What to consider when visualizing data for colorblind readers},\n year = {2020}\n}"
  ]
}
FieldAccepted value
idNon-empty string used by commands and APIs
titleString
descriptionString
labelsArray of canonical family:value strings
sectionsNon-empty array of valid role, title, and content objects
referencesArray of BibTeX strings cited by the sections

Readers join sections in array order to produce the guideline body. A citation such as [@muth_colorblindness_2020] refers to the matching BibTeX key in references.

Catalog locations

chartcoach opens these directory and file layouts:

LocationFilesBehavior
Authored folderMANIFEST.md and entries/<id>/guideline.mdReads the Markdown entries used for curation
Compiled bundleMANIFEST.md and entries.parquetReads local guideline entry rows directly
Local deployed rootcatalog.json and catalog/releases/Opens and verifies the selected release
Exact release directoryrelease.json plus the files it listsVerifies files as they are read
Remote selectionA catalog.json URLFollows the current release selection
Remote exact releaseA release.json URLKeeps one release digest

MANIFEST.md lists the section roles and label families accepted by that catalog. Readers reject records that use a role or label family missing from the manifest.

release.json

release.json lists each catalog and index file that readers may load. It records the byte count and SHA-256 hash that readers verify. Current readers accept schema_version: 1:

type CatalogRelease = {
  schema_version: 1;
  digest: string;
  artifacts: Record<
    string,
    {
      sha256: string;
      bytes: number;
    }
  >;
};

Every release includes MANIFEST.md and entries.parquet. Readers first check that digest matches the schema_version and path-sorted artifacts entries. They then verify each required file's byte count and hash before parsing it.

A release can also include a named index profile:

release/
├── release.json
├── MANIFEST.md
├── entries.parquet
└── profiles/
    └── minilm-normalized/
        ├── profile.json
        ├── index.tar.gz
        ├── documents.parquet    # optional
        └── projection.parquet   # optional

A profile ID is a lowercase portable single-component name scoped to one release. The embedding binding, dimensions, distance metric, Python requirements, and projection metadata belong in profile.json.

index.tar.gz contains the LanceDB table named documents, including document identities, text, labels, vectors, and embedding metadata. profile.json describes the index:

FieldMeaning
schema_version, documents_versionMetadata and document formats
entries_digest, manifest_digestCatalog contents indexed by this profile
embedding_functionsLanceDB registry bindings from text to vectors
dimensions, distance_metricVector size and retrieval distance
python_requirementsExact dependency versions recorded by the producer
lancedb_versionProducer's LanceDB version
projectionProjection algorithm and options, or null

The metadata file is at most 64 KiB and has a closed schema. Registry variable references use $var:name. Values remain caller-owned. Every profile contains this file and index.tar.gz.

documents.parquet is an optional portable export with row_id, id, parent_id, role, labels, content_hash, text, and vector. Validation compares it with the LanceDB table when present.

projection.parquet is present when profile.json.projection contains projection metadata. It has row_id, id, parent_id, role, projection_x, projection_y, and neighbors. Neighbor IDs refer to the shared row_id mapping. Distances come from the original vector-space neighbor graph used by the projection algorithm. They are not distances between points in the rendered two-dimensional projection.

catalog.json

catalog.json contains the record for the currently selected public release. Selecting another release overwrites this file. Files under catalog/releases/<digest>/ keep the bytes described by that digest.

Readers fetch catalog.json on each open so they see the current selection. Files already verified for the selected digest can come from cache. A specific release.json keeps the same digest across calls.

When the directory immediately before release.json is a SHA-256 digest, readers also require that path segment to match the descriptor digest.

Catalog tables

Python and the CLI derive six tables from the guideline entry records:

TableOne row per
guidelinesGuideline entry
sectionsGuideline section
guideline_labelsGuideline-label pair
referencesParsed BibTeX reference
guideline_referencesGuideline-reference pair
guideline_sourcesGuideline-reference pair with parsed source fields

Inspect the available tables, columns, and row counts before writing a query:

chartcoach catalog schema --tables --row-counts

On this page