Python
Open one Guideline Catalog and compose its entries with Polars, DuckDB, and LanceDB.
open_catalog() returns a Catalog for reading and citing guideline entries.
Use its native Polars dataframes for filtering,
DuckDB connections for
SQL, and LanceDB tables for search.
Install
Follow Getting started to install chartcoach.
Examples reuse the catalog opened in the first block
unless they explicitly open another source.
Open and read a catalog
from chartcoach import open_catalog
catalog = open_catalog()
record = catalog.read(
ids=["directly-label-series-instead-of-using-a-color-key"],
source_detail="minimal",
)[0]
print(record["title"])Expected output:
Directly label colored series instead of relying on a color keyopen_catalog(
location: str | PathLike[str] | None = None,
*,
storage_options: Mapping[str, object] | None = None,
) -> Cataloglocation accepts an authored folder, compiled bundle, local deployed root,
local release directory, or local or remote catalog.json or release.json.
A deployed root contains catalog.json and
catalog/releases/<digest>/. Omitting location opens the official selected
catalog. Local paths, file://, HTTP, and HTTPS use the base package. S3, GCS,
and Azure locations require chartcoach[cloud].
Release-backed opening verifies the descriptor digest plus the byte counts and
SHA-256 hashes of MANIFEST.md and entries.parquet. Profile files are
verified when an operation reads them. catalog.release contains the
immutable descriptor. Authored folders and compiled bundles set it to None.
Select, read, and cite guidelines
| Method | Returns |
|---|---|
query(*, ids=(), labels=(), label_prefixes=(), contains=None, limit=50) | Polars DataFrame of candidates |
read(*, ids, roles=(), source_detail="minimal") | List of guideline entry records |
cite(*, ids, url_template=...) | List of guideline and source citations |
candidates = catalog.query(contains="direct labels", limit=5)
ids = candidates.get_column("id").head(3).to_list()
records = catalog.read(ids=ids, source_detail="minimal")
citations = catalog.cite(ids=ids)catalog.query() returns a Polars DataFrame with id, title, description,
and labels. Exact guideline entry IDs retain their requested order. Repeated
labels are all-of filters. Substring matching ignores case and treats spaces,
hyphens, and underscores as equivalent separators. contains matches one
contiguous phrase in the ID, title, or description. Empty results retain the
candidate schema. Shorten the phrase, relax filters, or search section text
with SQL when needed.
catalog.read() returns guideline entry records in requested ID order,
including duplicates. roles filters section roles while retaining authored
order. source_detail accepts none, minimal, or full.
Full source detail returns every parsed source field and the entry's BibTeX
references.
catalog.cite() returns each guideline URL, its formatted citation, and its
source citations, formatted with RefKit's APA style and supplemented with source
locators. Pass url_template="https://catalog.example/g/{id}" for a
custom public catalog. A citation identifies an attached source. Inspect the
publication before attributing a specific claim to it.
read and cite require exact IDs and raise CatalogError with code lookup
for an unknown ID. Empty ID lists return empty lists.
Inspect catalog identity
description = catalog.describe()CatalogInfo contains resolved_location, entries_digest,
manifest_digest, release_digest, catalog table schemas and row counts,
vocabulary, profile names, and optional ProfileInfo. Credentials, URL query
parameters, and storage options stay private.
Pass profile=NAME to read that profile's verified profile.json. ProfileInfo
reports the profile and document schema versions, embedding binding,
dimensions, distance_metric, python_requirements, LanceDB version, and
projection metadata. Import CatalogInfo, TableInfo, TableColumnInfo, and ProfileInfo from chartcoach
when application code needs these annotations. The index stays unopened and
the embedding provider stays unconstructed.
| Member | Result |
|---|---|
len(catalog) | Number of guideline entries |
catalog.manifest | CatalogManifest with Markdown and vocabulary |
catalog.release | CatalogRelease, or None for an authored folder or bundle |
catalog.entries_digest() | SHA-256 digest of canonical entry records |
catalog.describe(*, profile=None) | Identity, schemas, row counts, vocabulary, and profiles |
Compose Polars tables
catalog.to_frame() returns the six canonical stored fields. The catalog
tables are:
catalog.table(name) | One row per |
|---|---|
guidelines | Guideline entry with derived Markdown body |
sections | Guideline section |
guideline_labels | Guideline-label pair |
references | Parsed BibTeX reference |
guideline_references | Guideline-reference pair |
guideline_sources | Guideline-reference pair with source fields |
Tables are derived on first use and reused within that Catalog. Every call
returns an isolated Polars DataFrame handle, so caller mutations leave the
catalog's cached tables unchanged.
import polars as pl
counts = (
catalog.table("guideline_labels")
.group_by("family")
.len()
.sort("len", descending=True)
)
line_guidelines = catalog.to_frame().filter(
pl.col("labels").list.contains("chart:line")
)For native index ingestion, catalog.documents() derives the canonical search
documents and their row_id mapping once per Catalog, returning an isolated
Polars DataFrame handle on each call. Add vectors using
LanceDB's embedding registry or the application's embedding pipeline, then use
chartcoach.curation.IndexProfile to package the table into a release.
Compose DuckDB queries
catalog.duckdb(config=None) returns a fresh caller-owned in-memory DuckDB
connection containing the six catalog tables. Each connection owns its tables,
so SQL mutations leave other connections and the catalog unchanged. Reuse a
connection for related queries and close it after use.
with catalog.duckdb() as connection:
rows = connection.sql("""
select g.id, g.title, s.role
from guidelines g
join sections s on s.guideline_id = g.id
order by g.id
limit 5
""").fetchall()Use the native connection for parameters, joins, aggregation, and exports.
connection.execute(statement, parameters).fetchall() binds values separately
from SQL. The CLI and MCP server provide bounded SQL results for their transports.
Use chartcoach.duckdb.register_catalog(connection, catalog, ids=None) to create
or replace the six tables in an existing connection. It returns the same
caller-owned connection and preserves unrelated tables. Pass ids=[...] to
register selected guideline entries and all their linked reference records.
Duplicate IDs are accepted. Omit ids for the full catalog, or pass ids=[]
for six empty typed tables. Unknown IDs and invalid references fail before
replacing tables.
write_duckdb(catalog, output, overwrite=False) writes a database file.
Open an index as a LanceDB table
Install the index capability in the active environment:
uv pip install "chartcoach[index]"This example requires an indexed release built with curation.
Use catalog.describe()["profiles"] to inspect the profiles in another release.
catalog = open_catalog("dist/release-indexed")
PROFILE = "minilm-normalized"
table = catalog.index(PROFILE)
fts = table.search(
"direct labels",
query_type="fts",
fts_columns="text",
).select(["id", "parent_id", "role", "_score"]).limit(5).to_list()The default table reads a protected shared extraction pinned to its published
version. Pass a new directory=Path(...) to create a caller-owned writable
extraction. The directory must not exist, and the caller owns its lifetime. A
platform without POSIX permission protection requires this explicit directory.
Numeric vectors, .select(), .where(), rerankers, batch queries, and index
tuning use the returned LanceDB table directly. Apply the distance_metric
reported by catalog.describe(profile=PROFILE) to vector queries.
The release builder creates a full-text index. Vector search scans stored
vectors until the release producer creates a LanceDB vector index explicitly.
Search and read matching guidelines
table = catalog.index(PROFILE)
hits = table.search(
"direct labels", query_type="fts", fts_columns="text"
).where("role = 'section.advice'").select(
["id", "parent_id", "role", "_score"]
).limit(10).to_list()
ids = list(dict.fromkeys(hit["parent_id"] for hit in hits))
records = catalog.read(ids=ids)
citations = catalog.cite(ids=ids)Explicit query_type="fts" keeps embedding providers idle. vector and hybrid
pass text to the embedding function persisted in the LanceDB table. Install
the exact python_requirements reported by
catalog.describe(profile=PROFILE) and register custom aliases before semantic
search.
For a profile whose binding uses $var:provider-key, load the secret from the
application's environment before text-based vector or hybrid search:
import os
from lancedb.embeddings import get_registry
get_registry().set_var("provider-key", os.environ["EMBEDDING_API_KEY"])
metadata = catalog.describe(profile=PROFILE)["profile"]
semantic = table.search("direct labels", query_type="vector").distance_type(
metadata["distance_metric"]
).select(["id", "parent_id", "role", "_distance"]).limit(10).to_list()Each index row has row_id, id, parent_id, role, labels,
content_hash, text, and vector. parent_id joins to the guideline entry
ID. Roles are overview, document, or section.<manifest-role>.
LanceDB limits document hits, so deduplicate parent_id in ranked order before
reading guidelines. An additional role = 'overview' query can broaden
candidates when many hits share one parent. Section queries recover details
that an overview may omit. Check the selected entries' context and exceptions
before applying their advice.
Project the columns needed for candidate selection before calling to_list().
Include _score for full-text queries and _distance for vector queries. For
hybrid queries, select stored fields such as id, parent_id, and role.
LanceDB adds _relevance_score after reranking. Higher relevance and lower
distance rank first. Scores compare retrieval results within a query, mode, and
profile. They do not measure confidence in a recommendation. Use native
selection, filtering, reranking, and batch APIs to compose other workflows.
Use verified files and reopen offline
entries = catalog.artifact("entries.parquet")
with catalog.duckdb() as connection:
rows = connection.read_parquet(str(entries)).select("id, title").limit(5).fetchall()
local = catalog.cache()
offline = open_catalog(local)catalog.artifact(path) accepts a path from catalog.release.artifacts and
returns a verified local file. Optional profile exports include
profiles/<profile>/documents.parquet and profiles/<profile>/projection.parquet.
Use the returned paths with DuckDB, Polars, or Arrow. An unknown artifact raises
CatalogError with code lookup.
Remote files use the per-user cache directory selected by
platformdirs. Cached bytes are checked
against their descriptor before reuse. Digest-addressed remote releases can
reopen from cached descriptors and core files. Opening a selection refreshes
catalog.json. catalog.cache() materializes every release artifact into a
local release directory, which can include large index and projection files.
Both methods require a release-backed catalog.
See Open and cache catalogs for cloud storage and credential setup.
Construct an in-memory catalog
With the MANIFEST.md from Author a guideline:
from chartcoach import Catalog, CatalogManifest, Guideline, Section
manifest = CatalogManifest.from_path("authored-catalog/MANIFEST.md")
guideline = Guideline(
"direct-labels", "Use direct labels", "Put labels close to marks.",
labels=("chart:line",),
sections=(Section("advice", "Advice", "Label each series directly."),),
)
catalog = Catalog.from_guidelines([guideline], manifest=manifest)The manifest must define advice and the chart label family. Use
Catalog(frame, manifest=manifest) for a Polars frame with the six stored fields.
Models and curation
Import these models from chartcoach. Use the parsing methods for external data:
| Model | Fields and operations |
|---|---|
Guideline | id, title, description, labels, sections, references, derived body, from_mapping(), to_record() |
Section | role, title, content, from_mapping(), to_record() |
CatalogManifest | markdown, section_roles, label_families, from_text(), from_path(), write() |
ManifestDefinition | Vocabulary name, description, and examples |
CatalogRelease | digest, artifacts, from_mapping(), to_record(), artifact(path) |
ReleaseArtifact | sha256, bytes, from_mapping(), to_record() |
Operation annotations are CatalogInfo, ProfileInfo, GuidelineEntryRecord,
SectionRecord, SourceDetail, MinimalSourceRecord, FullSourceRecord,
CitationRecord, and CitationSource. chartcoach.__version__ reports the
installed package version.
Import curation operations from chartcoach.curation:
| Operation | Result |
|---|---|
write_bundle(catalog, output) | Compiled catalog directory |
build_release(catalog, output, *, profiles={}) | CatalogRelease and its files |
validate_release(directory) | Validated CatalogRelease |
validate_published_release(destination, *, digest=None, storage_options=None) | Freshly validated published CatalogRelease |
publish_release(source, destination, *, storage_options=None) | Published CatalogRelease |
select_release(digest, destination, *, storage_options=None) | Selected CatalogRelease |
EmbeddingProfile, IndexProfile, and ProfileReuse describe index inputs.
ProfileBuild is their union. See Curate and publish for build
options, output-directory requirements, and publication side effects.
validate_published_release downloads fresh artifacts into temporary files.
With digest=None, it verifies catalog.json against the matching published
descriptor. Pass a digest to validate an exact candidate. Validation uses
stored vectors and leaves the destination and runtime cache unchanged.
Errors
Expected chartcoach failures raise CatalogError. Its code is one of
lookup, invalid_input, integrity, unavailable_capability,
incompatible_profile, embedding_failure, operation_failed, or
response_too_large.
details carries machine-readable recovery context, and hints describes the
next valid action. LanceDB objects retain their own exception behavior.