Blog

Articles on automated cataloging, catalog formats and use cases.

Why trust a catalog made with AI

How a catalog made with AI becomes trustworthy: not by denying that the model enriches, but by knowing where each piece of data comes from. It rests on the source that backs each one, verification against authorities, deterministic computation of rule-based fields, per-record evaluation and a quality gate that stops degraded batches, with the limits made explicit.

How to describe an archive in ISAD-G with AI

What sets archival description apart from bibliographic description, why an archive needs ISAD-G, and how ISAD-G description is produced with AI —by provenance and hierarchy, not item by isolated item—.

Glossary: the AI and cataloging terms in this blog

Short definitions of the terms that cross this series: from AI (language model or LLM, RAG, frontier model, on-premise, air-gapped, anonymization, OCR, ASR), from cataloging (MARC21, Dublin Core, ISAD-G, CDWA, authority control) and the Janium system. For readers coming from libraries and archives who run into AI jargon, or the other way around.

Cataloging with AI without the material leaving the institution

Some collections cannot leave the institution — data residency, policy or confidentiality. How Collect catalogs with AI on the institution's own infrastructure, with a local model and even air-gapped; the hybrid mode with anonymization for the hard cases; and why the rare thing isn't sovereign AI or cataloging on their own, but the two together. With the limits of each mode.

From documents to catalog records — what Janium Collect is

What problem Janium Collect solves for an institution with a collection, how it turns documents of different kinds into catalog records in standard formats, and where the limits of what it does lie.