Catalog data
Inspect guideline entries, catalog locations, and the files in a published release.
Each row in entries.parquet is one guideline entry. This
Parquet file stores tabular data for Python,
JavaScript, and database readers. Entries contain labels, ordered sections,
and BibTeX source records for citations.
Guideline entry record
This excerpt uses exact values from the quickstart's catalog. It shows the first section and complete reference:
{
"id": "directly-label-series-instead-of-using-a-color-key",
"title": "Directly label colored series instead of relying on a color key",
"description": "For color-coded series identification, prefer direct labels on chart marks to improve readability and mitigate color-key decoding for readers with color-vision deficiency.",
"labels": [
"purpose:refine",
"basis:heuristic",
"quality:accessibility",
"lever:text-annotation",
"component:label:use",
"component:legend:avoid",
"needs:color-vision-deficiency",
"polish:annotation"
],
"sections": [
{
"role": "advice",
"title": "Move labels onto the marks",
"content": "Label colored series directly on the chart instead of making readers decode them through a color key. For example, place series names next to lines or areas, or label pie slices directly so readers do not have to match colors back and forth with a legend."
}
],
"references": [
"@misc{muth_colorblindness_2020,\n author = {Lisa Charlotte Muth},\n howpublished = {\\url{https://www.datawrapper.de/blog/colorblindness-part2}},\n note = {Datawrapper Blog. Accessed: 2025-11-17},\n title = {What to consider when visualizing data for colorblind readers},\n year = {2020}\n}"
]
}| Field | Accepted value |
|---|---|
id | Non-empty string used by commands and APIs |
title | String |
description | String |
labels | Array of canonical family:value strings |
sections | Non-empty array of valid role, title, and content objects |
references | Array of BibTeX strings cited by the sections |
Readers join sections in array order to produce the guideline body. A
citation such as [@muth_colorblindness_2020] refers to the matching BibTeX
key in references.
Catalog locations
chartcoach opens these directory and file layouts:
| Location | Files | Behavior |
|---|---|---|
| Authored folder | MANIFEST.md and entries/<id>/guideline.md | Reads the Markdown entries used for curation |
| Compiled bundle | MANIFEST.md and entries.parquet | Reads local guideline entry rows directly |
| Local deployed root | catalog.json and catalog/releases/ | Opens and verifies the selected release |
| Exact release directory | release.json plus the files it lists | Verifies files as they are read |
| Remote selection | A catalog.json URL | Follows the current release selection |
| Remote exact release | A release.json URL | Keeps one release digest |
MANIFEST.md lists the section roles and label families accepted by that
catalog. Readers reject records that use a role or label family missing from
the manifest.
release.json
release.json lists each catalog and index file that readers may load.
It records the byte count and SHA-256 hash that readers verify. Current readers
accept schema_version: 1:
type CatalogRelease = {
schema_version: 1;
digest: string;
artifacts: Record<
string,
{
sha256: string;
bytes: number;
}
>;
};Every release includes MANIFEST.md and entries.parquet. Readers first
check that digest matches the schema_version and path-sorted artifacts
entries. They then verify each required file's byte count and hash before
parsing it.
A release can also include a named index profile:
release/
├── release.json
├── MANIFEST.md
├── entries.parquet
└── profiles/
└── minilm-normalized/
├── profile.json
├── index.tar.gz
├── documents.parquet # optional
└── projection.parquet # optionalA profile ID is a lowercase portable single-component name scoped to one
release. The embedding binding, dimensions, distance metric, Python
requirements, and projection metadata belong in profile.json.
index.tar.gz contains the LanceDB table named documents, including
document identities, text, labels, vectors, and embedding metadata.
profile.json describes the index:
| Field | Meaning |
|---|---|
schema_version, documents_version | Metadata and document formats |
entries_digest, manifest_digest | Catalog contents indexed by this profile |
embedding_functions | LanceDB registry bindings from text to vectors |
dimensions, distance_metric | Vector size and retrieval distance |
python_requirements | Exact dependency versions recorded by the producer |
lancedb_version | Producer's LanceDB version |
projection | Projection algorithm and options, or null |
The metadata file is at most 64 KiB and has a closed schema. Registry variable
references use $var:name. Values remain caller-owned. Every profile contains
this file and index.tar.gz.
documents.parquet is an optional portable export with row_id, id,
parent_id, role, labels, content_hash, text, and vector. Validation
compares it with the LanceDB table when present.
projection.parquet is present when profile.json.projection contains
projection metadata. It has row_id, id, parent_id, role,
projection_x, projection_y, and neighbors. Neighbor IDs refer to the
shared row_id mapping. Distances come from the original vector-space neighbor
graph used by the projection algorithm. They are not distances between points
in the rendered two-dimensional projection.
catalog.json
catalog.json contains the record for the currently selected public release.
Selecting another release overwrites this file. Files under
catalog/releases/<digest>/ keep the bytes described by that digest.
Readers fetch catalog.json on each open so they see the current selection.
Files already verified for the selected digest can come from cache. A specific
release.json keeps the same digest across calls.
When the directory immediately before release.json is a SHA-256 digest,
readers also require that path segment to match the descriptor digest.
Catalog tables
Python and the CLI derive six tables from the guideline entry records:
| Table | One row per |
|---|---|
guidelines | Guideline entry |
sections | Guideline section |
guideline_labels | Guideline-label pair |
references | Parsed BibTeX reference |
guideline_references | Guideline-reference pair |
guideline_sources | Guideline-reference pair with parsed source fields |
Inspect the available tables, columns, and row counts before writing a query:
chartcoach catalog schema --tables --row-counts