JavaScript and TypeScript
Open a catalog release, query guidance, read sources, and inspect profile metadata.
@chartcoach/catalog reads guideline entries and verified release files in
browsers and Node.js. Use @chartcoach/catalog/node for local paths and a
persistent file cache.
Follow Build a web app to see one catalog power browser-side exploration, server-side search, and guideline citations in the chat app.
Install
npm install @chartcoach/catalogUse Node.js 22.19 or later. Browser applications require the Web APIs and
WebAssembly setup described with openCatalog.
Save the first example as read-guideline.mjs in your project and run
node read-guideline.mjs.
Subsequent snippets assume the catalog created by this example unless they
open another catalog explicitly.
Open and read a catalog
import { openCatalog } from "@chartcoach/catalog";
const catalog = await openCatalog();
const candidates = catalog.query({ contains: "direct labels", limit: 5 });
const ids = candidates.map(({ id }) => id);
const records = catalog.read({ ids, sourceDetail: "minimal" });
const citations = catalog.cite({ ids });
const info = await catalog.describe();
console.log(records[0]?.title);The first line is:
Directly label colored series instead of relying on a color keyopenCatalog(location?, { fetch?, signal?, cache? }) accepts an absolute HTTP or HTTPS URL
whose path names catalog.json or release.json. Omit location to open the
official selected catalog. Pass a custom fetch implementation when the
application needs authentication or another URL scheme, such as S3. Pass an AbortSignal
to cancel opening. Later describe and artifact calls accept their own
signals. The runtime must provide Web Crypto,
AbortSignal.timeout, and AbortSignal.any.
Reference parsing and APA citations use RefKit.
Importing the SDK initializes its WebAssembly engine. Browser bundlers must emit
RefKit's .wasm asset at its module-relative URL. A restrictive
Content Security Policy
must allow 'wasm-unsafe-eval' in script-src and the asset origin in connect-src.
openCatalog verifies the release digest, then checks the byte count and
SHA-256 hash of MANIFEST.md and entries.parquet before parsing them.
catalog.json follows the current public selection. release.json keeps one
digest across calls. The returned catalog.release contains the verified
descriptor. catalog.releaseUrl contains the exact release.json location with
userinfo, query parameters, and fragments removed. Invalid locations, failed
requests, malformed release data, and integrity mismatches reject the promise
with CatalogError.
Load a compiled bundle
For a bundle built with curation, loadCatalogData accepts
entries.parquet bytes and MANIFEST.md text. The paths in this example are
relative to the process's working directory:
import { readFile } from "node:fs/promises";
import { loadCatalogData } from "@chartcoach/catalog";
const catalog = await loadCatalogData({
entries: await readFile("dist/catalog/entries.parquet"),
manifestText: await readFile("dist/catalog/MANIFEST.md", "utf8"),
});loadCatalogData validates the manifest vocabulary and guideline fields.
Release integrity remains caller-owned because this input contains no
release.json. Invalid manifests and decoded guideline entry rows cause
loadCatalogData to reject with CatalogError. Parquet decoding failures reject
the promise with the reader's error.
Load verified release files
loadCatalog verifies a parsed release and caller-provided core bytes. Use it
when your application already owns retrieval. This example requires a built
release at dist/release relative to the process's working directory:
import { readFile } from "node:fs/promises";
import { pathToFileURL } from "node:url";
import { loadCatalog, parseCatalogRelease } from "@chartcoach/catalog";
const releasePath = "dist/release/release.json";
const release = parseCatalogRelease(JSON.parse(await readFile(releasePath, "utf8")));
const catalog = await loadCatalog({
release,
releaseUrl: pathToFileURL(releasePath),
entries: await readFile("dist/release/entries.parquet"),
manifest: await readFile("dist/release/MANIFEST.md"),
});The returned catalog carries the copied release descriptor and sanitized exact
locator. A digest-addressed releaseUrl must agree with the descriptor digest.
Each core file is capped at 64 MiB. Use openCatalog when
describe({ profile }) should fetch profile.json.
Catalog methods
query, read, cite, and table return frozen JavaScript arrays synchronously from
the loaded Catalog. describe returns a promise because it hashes the loaded
records and can fetch profile metadata.
const candidates = catalog.query({
labels: ["purpose:refine"],
labelPrefixes: ["quality:"],
contains: "direct labels",
limit: 5,
});
const ids = candidates.map(({ id }) => id);
const records = catalog.read({ ids, roles: ["advice"], sourceDetail: "full" });
const citations = catalog.cite({ ids });
const description = await catalog.describe();query(options?)returns compactEntryCandidatevalues withid,title,description, andlabels. Labels and label prefixes are conjunctive. The default limit is 50.containsmatches a contiguous phrase in the ID, title, or description, ignoring case and normalizing whitespace, hyphens, and underscores. Empty matches return an empty array. Shorten the phrase or query section text through the application's data engine when needed.read({ ids, roles?, sourceDetail? })preserves requested guideline entry ID order, duplicate IDs, and authored section order. It returnsGuidelineEntryRecordvalues. Source detail defaults to"minimal".cite({ ids, urlTemplate? })returns guideline links and structured source citations. Equivalent BibTeX definitions share one source. Conflicting definitions for one key throwCatalogError.describe({ profile?, signal? })returns snake-case identity, vocabulary, profile names, and optionalProfileInfo. Calls with the default options are also awaited.
| Member | Result |
|---|---|
catalog.get(id) | Guideline, or undefined when absent |
catalog.require(id) | Guideline, or CatalogError with code lookup |
catalog.guidelines, iteration | Complete stored guidelines in catalog order |
catalog.length | Number of guideline entries |
catalog.labels() | Sorted distinct labels |
catalog.sectionRoles() | Sorted distinct roles used by guidelines |
catalog.manifest | Manifest Markdown and vocabulary definitions |
catalog.release, catalog.releaseUrl | Release descriptor and sanitized exact location, when present |
toMarkdown(guideline) | Authored Markdown with frontmatter and ordered sections |
JavaScript options use camel case. Shared operation records use snake case for
fields such as source_title, reference_id, and release_digest. Nullable
JSON fields are null. The top-level references field appears on full-detail
read records. Full parsed source objects omit BibTeX because references is
the raw-reference authority for each entry.
Guideline values, operation results, their nested arrays, and manifest
definitions are frozen. Pass changed record objects to a new Catalog. Use
loadCatalogData after writing changed rows to entries.parquet.
Catalogs created with loadCatalogData expose undefined for release and
releaseUrl because that call receives the two bundle files directly.
CatalogError exposes code, details, and hints for programmatic recovery.
Compose catalog tables
catalog.table(name) returns cached immutable rows. Its six table names and
fields match Python's catalog.table(name):
| Table | Rows |
|---|---|
guidelines | Guideline entries with derived Markdown and nested sections |
sections | Individual sections linked by guideline_id |
guideline_labels | Parsed labels linked by guideline_id |
references | Distinct bibliographic sources, including authored BibTeX |
guideline_references | Guideline-to-reference links |
guideline_sources | Linked source details for each guideline |
const sources = catalog.table("guideline_sources");
console.log(sources[0]?.authors);
const { tables } = await catalog.describe();
console.log(tables.map(({ name, rows }) => ({ name, rows })));tables contains { name, rows, columns }. Columns use the same logical type
names in Python and JavaScript, such as String and List(String).
Schema metadata remains available for empty tables. Unknown table names raise
CatalogError with code lookup. Parsed references and table projections are
cached per catalog. Nested arrays and records are frozen.
Register tables in DuckDB
Install DuckDB's Node.js client to use the optional adapter:
npm install @chartcoach/catalog @duckdb/node-apiimport { DuckDBInstance } from "@duckdb/node-api";
import { registerCatalog } from "@chartcoach/catalog/duckdb";
const db = await DuckDBInstance.create(":memory:");
const connection = await db.connect();
try {
await registerCatalog(connection, catalog);
const result = await connection.runAndReadAll(
'SELECT source_type, count(*) FROM "references" GROUP BY source_type',
);
console.log(result.getRowsJson());
} finally {
connection.closeSync();
db.closeSync();
}registerCatalog(connection, catalog, options?) returns a promise for the same
native connection. It creates or replaces the six named catalog tables and
preserves unrelated tables. The caller owns configuration, transactions, and
connection lifetime.
Pass { ids: [...] } to register selected guideline entries and all their linked
references. Duplicate IDs are accepted. Omitted ids selects every entry, while
{ ids: [] } creates six empty tables with complete schemas. Unknown IDs and
invalid references fail before replacing tables.
The adapter exports RegisterCatalogOptions from @chartcoach/catalog/duckdb.
DuckDB remains an optional peer dependency. Its native SQL results report native
types such as VARCHAR, independently of the catalog's logical schema metadata.
Query in a browser worker
@duckdb/duckdb-wasm runs
DuckDB in a browser worker. Install it alongside the catalog package:
npm install @chartcoach/catalog @duckdb/duckdb-wasmCreate and own the worker and connection using
DuckDB's browser setup.
Given that native AsyncDuckDBConnection as connection and the loaded catalog:
import { registerCatalog } from "@chartcoach/catalog/duckdb-wasm";
await registerCatalog(connection, catalog);
const result = await connection.query(
'SELECT source_type, count(*) FROM "references" GROUP BY source_type',
);
console.log(result.toArray());registerCatalog(connection, catalog, { ids? }) returns the same connection and
has the same table, selection, and transaction contract as the Node adapter.
RegisterCatalogOptions is exported from this entry point. Workers, connections,
and database instances remain caller-owned. Registration releases its temporary
in-memory files after materializing the tables.
For native Parquet queries, register the verified artifact bytes with DuckDB's filesystem. The catalog reuses bytes fetched while opening the release:
await connection.bindings.registerFileBuffer(
"entries.parquet",
await catalog.artifact("entries.parquet"),
);
try {
const result = await connection.query(
"SELECT id, title, labels FROM read_parquet('entries.parquet') LIMIT 5",
);
console.log(result.toArray());
} finally {
await connection.bindings.dropFile("entries.parquet");
}Use DuckDB's EH bundle with a browser that supports WebAssembly exception handling.
The chat app checks that capability and serves the worker and WASM files from
the installed package. A restrictive
Content Security Policy must allow workers, 'wasm-unsafe-eval', and the
configured DuckDB extension repository in connect-src. The default repository
is https://extensions.duckdb.org, which supplies the version-matched JSON
extension on first use. TypeScript projects using the current DuckDB-WASM
declarations may need skipLibCheck for its upstream declaration dependencies.
The Build a web app guide demonstrates this setup with Mosaic, a library that coordinates SQL queries and linked selections across interactive views.
Inspect profile metadata
const info = await catalog.describe();
const profile = info.profiles[0];
if (profile !== undefined) {
const controller = new AbortController();
const details = await catalog.describe({ profile, signal: controller.signal });
console.log(details.profile?.distance_metric, details.profile?.dimensions);
}The profile call verifies the profile.json byte count, SHA-256 hash, 64 KiB
bound, entries digest, and manifest digest. It reads metadata from the release
resolved by openCatalog and keeps the index unopened. Successful metadata is
cached per Catalog instance. Cancellation affects one description call.
ProfileInfo contains the profile and document schema versions, embedding
binding, dimensions, distance_metric, python_requirements, LanceDB version,
and projection metadata.
parseProfileMetadata(value) validates an already parsed profile.json value
and returns immutable metadata. It checks exact fields, digests, dimensions,
distance metric, LanceDB embedding binding, exact Python requirement versions,
projection metadata, and portable registry-variable syntax. Profile IDs use
lowercase single-component names such as minilm-normalized.
Parsing validates the record shape. Use catalog.describe({ profile }) to
also verify the metadata's bytes and relationship to the loaded catalog.
Use release files with native tools
const bytes = await catalog.artifact("entries.parquet");
const artifacts = Object.keys(catalog.release?.artifacts ?? {});artifact(path, { signal? }) returns verified Uint8Array bytes. The release
inventory includes available document and projection Parquet exports beneath
profiles/<profile>/. Core files and small metadata are reused in memory.
For application-owned persistence, pass { cache } to openCatalog. An
ArtifactCache implements:
interface ArtifactCache {
get(sha256: string, size: number): Promise<Uint8Array | undefined>;
put(sha256: string, bytes: Uint8Array): Promise<void>;
}Cached bytes are verified before reuse. Core files are limited to 64 MiB each. Optional artifact byte reads are limited to 1 GiB.
Node.js files and cache
For a built release at dist/release, request native file paths:
import { openCatalog, artifactPath } from "@chartcoach/catalog/node";
const catalog = await openCatalog("./dist/release");
const path = await artifactPath(catalog, "entries.parquet");Pass path to the DuckDB Node client
or an Arrow/Parquet reader. artifactPath verifies cached files and streams
missing files to disk. It accepts an optional third argument { signal }.
The default cache follows the platform's per-user cache location and shares
Python's chartcoach artifact layout. Pass { cacheDirectory: "/data/catalog-cache" }
to Node's openCatalog to choose another directory.
The Node entry accepts bundle directories, local release directories, deployed roots, local descriptor paths, and remote descriptor URLs. Exact digest-addressed releases reopen offline once their core files are cached. Selection descriptors refresh on each open. Optional files are cached on demand.
Query an index with LanceDB
LanceDB opens the extracted search database and provides native full-text and vector queries. Install its client in the application:
npm install @lancedb/lancedbOpen a release built with an index profile. This example uses ./dist/release:
import { connect } from "@lancedb/lancedb";
import { openCatalog, indexPath } from "@chartcoach/catalog/node";
const catalog = await openCatalog("./dist/release");
const [profile] = (await catalog.describe()).profiles;
if (!profile) throw new Error("Choose a catalog release with an index profile.");
const db = await connect(await indexPath(catalog, profile));
const table = await db.openTable("documents");
try {
await table.checkout(await table.version());
const matches = await table.query().fullTextSearch("direct labels").limit(5).toArray();
const ids = [...new Set(matches.map((row) => row.parent_id))];
console.log(catalog.read({ ids }));
} finally {
table.close();
db.close();
}indexPath(catalog, profile, options?)
Returns a promise for the extracted database directory containing
documents.lance. The catalog must come from a release whose profile metadata
can be loaded. The function verifies the profile's relationship to the catalog,
checks the archive's size and SHA-256 hash, and extracts bounded regular files.
Unsafe paths, links, and colliding names reject the operation.
catalog: A catalog returned byopenCatalog.profile: A name listed bycatalog.describe().options.directory: A new directory for a writable, caller-owned extraction. Existing directories are rejected. The caller owns its cleanup.options.signal: Cancels metadata loading, artifact retrieval, or extraction.
On POSIX systems, omitting directory reuses a completed read-only cache
generation. Concurrent callers do not receive partially extracted data. Windows
requires directory because POSIX permission protection is unavailable there.
Failed or cancelled extraction cleans the incomplete directory.
Pin the native table with checkout before reading shared cached data. To modify
an index, request a caller-owned directory and open that directory with LanceDB.
Use table.vectorSearch(vector) for query vectors produced by your application.
The profile's distance metric and vector dimensions are available through
catalog.describe({ profile }).
Full-text and numeric-vector queries require no embedding provider credentials. For string-based embedding queries, configure a compatible function through LanceDB's JavaScript embedding registry. A stored Python embedding binding does not register a JavaScript provider.
Open S3 with an existing storage client
An application can supply its own AWS SDK S3 client.
Install @aws-sdk/client-s3 in that application and configure its credentials
through the SDK's credential providers. Set CATALOG_SOURCE to your S3
descriptor URI. The adapter returns the SDK response body as a web stream and
forwards cancellation:
import { GetObjectCommand, S3Client } from "@aws-sdk/client-s3";
import { openCatalog } from "@chartcoach/catalog/node";
const s3 = new S3Client({ region: "eu-central-1" });
const catalog = await openCatalog(process.env.CATALOG_SOURCE, {
fetch: async (input, init) => {
const url = new URL(input);
const object = await s3.send(
new GetObjectCommand({
Bucket: url.hostname,
Key: decodeURIComponent(url.pathname.slice(1)),
}),
{ abortSignal: init?.signal ?? undefined },
);
if (!object.Body) throw new Error("S3 object body is missing");
return new Response(object.Body.transformToWebStream());
},
});chartcoach verifies descriptor and artifact bytes independently of transport. Browser consumers use the same custom-fetch option with the browser entry point and a credential strategy appropriate for the application.
Construct a catalog from records
Given manifest Markdown in manifestText and entry records in guidelines:
import { Catalog, parseCatalogManifest } from "@chartcoach/catalog";
const manifest = parseCatalogManifest(manifestText);
const catalog = new Catalog(guidelines, manifest);guidelines is an iterable of GuidelineInput records. Each contains id,
title, description, labels, sections, and BibTeX references. Each
GuidelineSection contains role, title, and content. The constructor
derives the Markdown body and validates vocabulary coverage.
Types and parsers
Import these contracts from @chartcoach/catalog:
| Group | Exports |
|---|---|
| Entries | GuidelineInput, Guideline, GuidelineSection, EntryCandidate, GuidelineEntryRecord |
| Sources | SourceDetail, MinimalSourceRecord, FullSourceRecord, CitationRecord, CitationSource |
| Catalog description | CatalogInfo, ProfileInfo |
| Tables | CatalogTables, TableName, TableInfo, TableColumnInfo |
| Manifest | CatalogManifest, ManifestDefinition, parseCatalogManifest |
| Release | CatalogRelease, ReleaseArtifact, parseCatalogRelease |
| Profiles | ProfileMetadata, EmbeddingBinding, ProjectionMetadata, DistanceMetric, parseProfileMetadata |
| Options | OpenCatalogOptions, LoadCatalogInput, LoadCatalogDataInput, QueryOptions, ReadOptions, CiteOptions, DescribeOptions, ArtifactOptions |
| Transport | FetchLike, ArtifactCache |
| Errors and JSON | CatalogError, CatalogErrorCode, CatalogErrorOptions, JsonObject, JsonValue |
parseCatalogManifest accepts Markdown text. parseCatalogRelease and
parseProfileMetadata accept parsed JSON values and validate their fields.
Loaders additionally verify release integrity. The Node entry exports
openCatalog, artifactPath, indexPath, NodeOpenCatalogOptions, and
IndexPathOptions.