chartcoach
Reference

JavaScript and TypeScript

Open a catalog release, query guidance, read sources, and inspect profile metadata.

@chartcoach/catalog reads guideline entries and verified release files in browsers and Node.js. Use @chartcoach/catalog/node for local paths and a persistent file cache.

Follow Build a web app to see one catalog power browser-side exploration, server-side search, and guideline citations in the chat app.

Install

npm install @chartcoach/catalog

Use Node.js 22.19 or later. Browser applications require the Web APIs and WebAssembly setup described with openCatalog.

Save the first example as read-guideline.mjs in your project and run node read-guideline.mjs. Subsequent snippets assume the catalog created by this example unless they open another catalog explicitly.

Open and read a catalog

read-guideline.mjs
import { openCatalog } from "@chartcoach/catalog";

const catalog = await openCatalog();
const candidates = catalog.query({ contains: "direct labels", limit: 5 });
const ids = candidates.map(({ id }) => id);
const records = catalog.read({ ids, sourceDetail: "minimal" });
const citations = catalog.cite({ ids });
const info = await catalog.describe();

console.log(records[0]?.title);

The first line is:

Directly label colored series instead of relying on a color key

openCatalog(location?, { fetch?, signal?, cache? }) accepts an absolute HTTP or HTTPS URL whose path names catalog.json or release.json. Omit location to open the official selected catalog. Pass a custom fetch implementation when the application needs authentication or another URL scheme, such as S3. Pass an AbortSignal to cancel opening. Later describe and artifact calls accept their own signals. The runtime must provide Web Crypto, AbortSignal.timeout, and AbortSignal.any.

Reference parsing and APA citations use RefKit. Importing the SDK initializes its WebAssembly engine. Browser bundlers must emit RefKit's .wasm asset at its module-relative URL. A restrictive Content Security Policy must allow 'wasm-unsafe-eval' in script-src and the asset origin in connect-src.

openCatalog verifies the release digest, then checks the byte count and SHA-256 hash of MANIFEST.md and entries.parquet before parsing them. catalog.json follows the current public selection. release.json keeps one digest across calls. The returned catalog.release contains the verified descriptor. catalog.releaseUrl contains the exact release.json location with userinfo, query parameters, and fragments removed. Invalid locations, failed requests, malformed release data, and integrity mismatches reject the promise with CatalogError.

Load a compiled bundle

For a bundle built with curation, loadCatalogData accepts entries.parquet bytes and MANIFEST.md text. The paths in this example are relative to the process's working directory:

import { readFile } from "node:fs/promises";
import { loadCatalogData } from "@chartcoach/catalog";

const catalog = await loadCatalogData({
  entries: await readFile("dist/catalog/entries.parquet"),
  manifestText: await readFile("dist/catalog/MANIFEST.md", "utf8"),
});

loadCatalogData validates the manifest vocabulary and guideline fields. Release integrity remains caller-owned because this input contains no release.json. Invalid manifests and decoded guideline entry rows cause loadCatalogData to reject with CatalogError. Parquet decoding failures reject the promise with the reader's error.

Load verified release files

loadCatalog verifies a parsed release and caller-provided core bytes. Use it when your application already owns retrieval. This example requires a built release at dist/release relative to the process's working directory:

import { readFile } from "node:fs/promises";
import { pathToFileURL } from "node:url";
import { loadCatalog, parseCatalogRelease } from "@chartcoach/catalog";

const releasePath = "dist/release/release.json";
const release = parseCatalogRelease(JSON.parse(await readFile(releasePath, "utf8")));
const catalog = await loadCatalog({
  release,
  releaseUrl: pathToFileURL(releasePath),
  entries: await readFile("dist/release/entries.parquet"),
  manifest: await readFile("dist/release/MANIFEST.md"),
});

The returned catalog carries the copied release descriptor and sanitized exact locator. A digest-addressed releaseUrl must agree with the descriptor digest. Each core file is capped at 64 MiB. Use openCatalog when describe({ profile }) should fetch profile.json.

Catalog methods

query, read, cite, and table return frozen JavaScript arrays synchronously from the loaded Catalog. describe returns a promise because it hashes the loaded records and can fetch profile metadata.

const candidates = catalog.query({
  labels: ["purpose:refine"],
  labelPrefixes: ["quality:"],
  contains: "direct labels",
  limit: 5,
});

const ids = candidates.map(({ id }) => id);
const records = catalog.read({ ids, roles: ["advice"], sourceDetail: "full" });
const citations = catalog.cite({ ids });
const description = await catalog.describe();
  • query(options?) returns compact EntryCandidate values with id, title, description, and labels. Labels and label prefixes are conjunctive. The default limit is 50. contains matches a contiguous phrase in the ID, title, or description, ignoring case and normalizing whitespace, hyphens, and underscores. Empty matches return an empty array. Shorten the phrase or query section text through the application's data engine when needed.
  • read({ ids, roles?, sourceDetail? }) preserves requested guideline entry ID order, duplicate IDs, and authored section order. It returns GuidelineEntryRecord values. Source detail defaults to "minimal".
  • cite({ ids, urlTemplate? }) returns guideline links and structured source citations. Equivalent BibTeX definitions share one source. Conflicting definitions for one key throw CatalogError.
  • describe({ profile?, signal? }) returns snake-case identity, vocabulary, profile names, and optional ProfileInfo. Calls with the default options are also awaited.
MemberResult
catalog.get(id)Guideline, or undefined when absent
catalog.require(id)Guideline, or CatalogError with code lookup
catalog.guidelines, iterationComplete stored guidelines in catalog order
catalog.lengthNumber of guideline entries
catalog.labels()Sorted distinct labels
catalog.sectionRoles()Sorted distinct roles used by guidelines
catalog.manifestManifest Markdown and vocabulary definitions
catalog.release, catalog.releaseUrlRelease descriptor and sanitized exact location, when present
toMarkdown(guideline)Authored Markdown with frontmatter and ordered sections

JavaScript options use camel case. Shared operation records use snake case for fields such as source_title, reference_id, and release_digest. Nullable JSON fields are null. The top-level references field appears on full-detail read records. Full parsed source objects omit BibTeX because references is the raw-reference authority for each entry.

Guideline values, operation results, their nested arrays, and manifest definitions are frozen. Pass changed record objects to a new Catalog. Use loadCatalogData after writing changed rows to entries.parquet. Catalogs created with loadCatalogData expose undefined for release and releaseUrl because that call receives the two bundle files directly.

CatalogError exposes code, details, and hints for programmatic recovery.

Compose catalog tables

catalog.table(name) returns cached immutable rows. Its six table names and fields match Python's catalog.table(name):

TableRows
guidelinesGuideline entries with derived Markdown and nested sections
sectionsIndividual sections linked by guideline_id
guideline_labelsParsed labels linked by guideline_id
referencesDistinct bibliographic sources, including authored BibTeX
guideline_referencesGuideline-to-reference links
guideline_sourcesLinked source details for each guideline
const sources = catalog.table("guideline_sources");
console.log(sources[0]?.authors);
const { tables } = await catalog.describe();
console.log(tables.map(({ name, rows }) => ({ name, rows })));

tables contains { name, rows, columns }. Columns use the same logical type names in Python and JavaScript, such as String and List(String). Schema metadata remains available for empty tables. Unknown table names raise CatalogError with code lookup. Parsed references and table projections are cached per catalog. Nested arrays and records are frozen.

Register tables in DuckDB

Install DuckDB's Node.js client to use the optional adapter:

npm install @chartcoach/catalog @duckdb/node-api
import { DuckDBInstance } from "@duckdb/node-api";
import { registerCatalog } from "@chartcoach/catalog/duckdb";

const db = await DuckDBInstance.create(":memory:");
const connection = await db.connect();
try {
  await registerCatalog(connection, catalog);
  const result = await connection.runAndReadAll(
    'SELECT source_type, count(*) FROM "references" GROUP BY source_type',
  );
  console.log(result.getRowsJson());
} finally {
  connection.closeSync();
  db.closeSync();
}

registerCatalog(connection, catalog, options?) returns a promise for the same native connection. It creates or replaces the six named catalog tables and preserves unrelated tables. The caller owns configuration, transactions, and connection lifetime.

Pass { ids: [...] } to register selected guideline entries and all their linked references. Duplicate IDs are accepted. Omitted ids selects every entry, while { ids: [] } creates six empty tables with complete schemas. Unknown IDs and invalid references fail before replacing tables.

The adapter exports RegisterCatalogOptions from @chartcoach/catalog/duckdb. DuckDB remains an optional peer dependency. Its native SQL results report native types such as VARCHAR, independently of the catalog's logical schema metadata.

Query in a browser worker

@duckdb/duckdb-wasm runs DuckDB in a browser worker. Install it alongside the catalog package:

npm install @chartcoach/catalog @duckdb/duckdb-wasm

Create and own the worker and connection using DuckDB's browser setup. Given that native AsyncDuckDBConnection as connection and the loaded catalog:

import { registerCatalog } from "@chartcoach/catalog/duckdb-wasm";

await registerCatalog(connection, catalog);
const result = await connection.query(
  'SELECT source_type, count(*) FROM "references" GROUP BY source_type',
);
console.log(result.toArray());

registerCatalog(connection, catalog, { ids? }) returns the same connection and has the same table, selection, and transaction contract as the Node adapter. RegisterCatalogOptions is exported from this entry point. Workers, connections, and database instances remain caller-owned. Registration releases its temporary in-memory files after materializing the tables.

For native Parquet queries, register the verified artifact bytes with DuckDB's filesystem. The catalog reuses bytes fetched while opening the release:

await connection.bindings.registerFileBuffer(
  "entries.parquet",
  await catalog.artifact("entries.parquet"),
);
try {
  const result = await connection.query(
    "SELECT id, title, labels FROM read_parquet('entries.parquet') LIMIT 5",
  );
  console.log(result.toArray());
} finally {
  await connection.bindings.dropFile("entries.parquet");
}

Use DuckDB's EH bundle with a browser that supports WebAssembly exception handling. The chat app checks that capability and serves the worker and WASM files from the installed package. A restrictive Content Security Policy must allow workers, 'wasm-unsafe-eval', and the configured DuckDB extension repository in connect-src. The default repository is https://extensions.duckdb.org, which supplies the version-matched JSON extension on first use. TypeScript projects using the current DuckDB-WASM declarations may need skipLibCheck for its upstream declaration dependencies.

The Build a web app guide demonstrates this setup with Mosaic, a library that coordinates SQL queries and linked selections across interactive views.

Inspect profile metadata

const info = await catalog.describe();
const profile = info.profiles[0];

if (profile !== undefined) {
  const controller = new AbortController();
  const details = await catalog.describe({ profile, signal: controller.signal });
  console.log(details.profile?.distance_metric, details.profile?.dimensions);
}

The profile call verifies the profile.json byte count, SHA-256 hash, 64 KiB bound, entries digest, and manifest digest. It reads metadata from the release resolved by openCatalog and keeps the index unopened. Successful metadata is cached per Catalog instance. Cancellation affects one description call.

ProfileInfo contains the profile and document schema versions, embedding binding, dimensions, distance_metric, python_requirements, LanceDB version, and projection metadata.

parseProfileMetadata(value) validates an already parsed profile.json value and returns immutable metadata. It checks exact fields, digests, dimensions, distance metric, LanceDB embedding binding, exact Python requirement versions, projection metadata, and portable registry-variable syntax. Profile IDs use lowercase single-component names such as minilm-normalized.

Parsing validates the record shape. Use catalog.describe({ profile }) to also verify the metadata's bytes and relationship to the loaded catalog.

Use release files with native tools

const bytes = await catalog.artifact("entries.parquet");
const artifacts = Object.keys(catalog.release?.artifacts ?? {});

artifact(path, { signal? }) returns verified Uint8Array bytes. The release inventory includes available document and projection Parquet exports beneath profiles/<profile>/. Core files and small metadata are reused in memory. For application-owned persistence, pass { cache } to openCatalog. An ArtifactCache implements:

interface ArtifactCache {
  get(sha256: string, size: number): Promise<Uint8Array | undefined>;
  put(sha256: string, bytes: Uint8Array): Promise<void>;
}

Cached bytes are verified before reuse. Core files are limited to 64 MiB each. Optional artifact byte reads are limited to 1 GiB.

Node.js files and cache

For a built release at dist/release, request native file paths:

import { openCatalog, artifactPath } from "@chartcoach/catalog/node";

const catalog = await openCatalog("./dist/release");
const path = await artifactPath(catalog, "entries.parquet");

Pass path to the DuckDB Node client or an Arrow/Parquet reader. artifactPath verifies cached files and streams missing files to disk. It accepts an optional third argument { signal }. The default cache follows the platform's per-user cache location and shares Python's chartcoach artifact layout. Pass { cacheDirectory: "/data/catalog-cache" } to Node's openCatalog to choose another directory.

The Node entry accepts bundle directories, local release directories, deployed roots, local descriptor paths, and remote descriptor URLs. Exact digest-addressed releases reopen offline once their core files are cached. Selection descriptors refresh on each open. Optional files are cached on demand.

Query an index with LanceDB

LanceDB opens the extracted search database and provides native full-text and vector queries. Install its client in the application:

npm install @lancedb/lancedb

Open a release built with an index profile. This example uses ./dist/release:

import { connect } from "@lancedb/lancedb";
import { openCatalog, indexPath } from "@chartcoach/catalog/node";

const catalog = await openCatalog("./dist/release");
const [profile] = (await catalog.describe()).profiles;
if (!profile) throw new Error("Choose a catalog release with an index profile.");

const db = await connect(await indexPath(catalog, profile));
const table = await db.openTable("documents");
try {
  await table.checkout(await table.version());
  const matches = await table.query().fullTextSearch("direct labels").limit(5).toArray();
  const ids = [...new Set(matches.map((row) => row.parent_id))];
  console.log(catalog.read({ ids }));
} finally {
  table.close();
  db.close();
}

indexPath(catalog, profile, options?)

Returns a promise for the extracted database directory containing documents.lance. The catalog must come from a release whose profile metadata can be loaded. The function verifies the profile's relationship to the catalog, checks the archive's size and SHA-256 hash, and extracts bounded regular files. Unsafe paths, links, and colliding names reject the operation.

  • catalog: A catalog returned by openCatalog.
  • profile: A name listed by catalog.describe().
  • options.directory: A new directory for a writable, caller-owned extraction. Existing directories are rejected. The caller owns its cleanup.
  • options.signal: Cancels metadata loading, artifact retrieval, or extraction.

On POSIX systems, omitting directory reuses a completed read-only cache generation. Concurrent callers do not receive partially extracted data. Windows requires directory because POSIX permission protection is unavailable there. Failed or cancelled extraction cleans the incomplete directory.

Pin the native table with checkout before reading shared cached data. To modify an index, request a caller-owned directory and open that directory with LanceDB. Use table.vectorSearch(vector) for query vectors produced by your application. The profile's distance metric and vector dimensions are available through catalog.describe({ profile }).

Full-text and numeric-vector queries require no embedding provider credentials. For string-based embedding queries, configure a compatible function through LanceDB's JavaScript embedding registry. A stored Python embedding binding does not register a JavaScript provider.

Open S3 with an existing storage client

An application can supply its own AWS SDK S3 client. Install @aws-sdk/client-s3 in that application and configure its credentials through the SDK's credential providers. Set CATALOG_SOURCE to your S3 descriptor URI. The adapter returns the SDK response body as a web stream and forwards cancellation:

import { GetObjectCommand, S3Client } from "@aws-sdk/client-s3";
import { openCatalog } from "@chartcoach/catalog/node";

const s3 = new S3Client({ region: "eu-central-1" });
const catalog = await openCatalog(process.env.CATALOG_SOURCE, {
  fetch: async (input, init) => {
    const url = new URL(input);
    const object = await s3.send(
      new GetObjectCommand({
        Bucket: url.hostname,
        Key: decodeURIComponent(url.pathname.slice(1)),
      }),
      { abortSignal: init?.signal ?? undefined },
    );
    if (!object.Body) throw new Error("S3 object body is missing");
    return new Response(object.Body.transformToWebStream());
  },
});

chartcoach verifies descriptor and artifact bytes independently of transport. Browser consumers use the same custom-fetch option with the browser entry point and a credential strategy appropriate for the application.

Construct a catalog from records

Given manifest Markdown in manifestText and entry records in guidelines:

import { Catalog, parseCatalogManifest } from "@chartcoach/catalog";

const manifest = parseCatalogManifest(manifestText);
const catalog = new Catalog(guidelines, manifest);

guidelines is an iterable of GuidelineInput records. Each contains id, title, description, labels, sections, and BibTeX references. Each GuidelineSection contains role, title, and content. The constructor derives the Markdown body and validates vocabulary coverage.

Types and parsers

Import these contracts from @chartcoach/catalog:

GroupExports
EntriesGuidelineInput, Guideline, GuidelineSection, EntryCandidate, GuidelineEntryRecord
SourcesSourceDetail, MinimalSourceRecord, FullSourceRecord, CitationRecord, CitationSource
Catalog descriptionCatalogInfo, ProfileInfo
TablesCatalogTables, TableName, TableInfo, TableColumnInfo
ManifestCatalogManifest, ManifestDefinition, parseCatalogManifest
ReleaseCatalogRelease, ReleaseArtifact, parseCatalogRelease
ProfilesProfileMetadata, EmbeddingBinding, ProjectionMetadata, DistanceMetric, parseProfileMetadata
OptionsOpenCatalogOptions, LoadCatalogInput, LoadCatalogDataInput, QueryOptions, ReadOptions, CiteOptions, DescribeOptions, ArtifactOptions
TransportFetchLike, ArtifactCache
Errors and JSONCatalogError, CatalogErrorCode, CatalogErrorOptions, JsonObject, JsonValue

parseCatalogManifest accepts Markdown text. parseCatalogRelease and parseProfileMetadata accept parsed JSON values and validate their fields. Loaders additionally verify release integrity. The Node entry exports openCatalog, artifactPath, indexPath, NodeOpenCatalogOptions, and IndexPathOptions.

On this page