Skip to main content
Documents are the raw units of knowledge: files, pages, extracts. nora documents covers the full library CRUD surface: uploads (single file, large file, whole folder), external-URL registration, metadata patches (regulatory dates for time-aware retrieval), chunk inspection, tagging, supersede, and delete.

Command groups

  • Read: list, get, chunks, chunk-text, extracted-text
  • Ingest: upload, upload-large, upload-folder, create-external
  • Mutate: tag, update, supersede, delete
  • KG grafting: extract-structure

Read

list

One row per document: id, title, content type, size, tags, last update. --tag narrows to one library folder. --json for scriptable output.

get

Full metadata: title, source URL, content hash, tags, ingest pipeline, chunk count, supersede state.

chunks

Lists a document’s persisted chunks. --preview caps at ~200 chars (or --preview-len <N>). --q runs a free-text filter inside chunk text.

chunk-text

Fetches one chunk’s full text, bypassing the preview cap.

extracted-text

The document’s cached extracted text (OCR / PDF / text extractor output). Returns null when extraction hasn’t run yet.

Ingest

upload: inline

Base64-inlines the file to the server. Cap: ~45 MB original / ~60 MB post-encoded. --mime optional (defaults to application/octet-stream). --name optional (defaults to the file’s basename). --tags stamps folder tags on the row.

upload-large: S3 presigned

Asks the server for an S3 presigned PUT URL, then PUTs the bytes directly, with no base64 overhead. Requires the deployment to have NORA_S3_BUCKET configured (otherwise a clear error). Use this above ~45 MB.

upload-folder: batch

Walks a folder and uploads every file. Auto-picks the transport per file: base64 inline under ~45 MB, S3-presigned above. --presigned forces the S3 path (useful when the tenant restricts inline payloads or the network prefers direct S3 PUT). Prints a progress line per file and a summary at the end. Idempotent per file: rerun to catch up after a mid-batch failure.

create-external: reference an URL

Registers an external URL as a document row (storage_kind='external'). Nora references the URL. It does not copy the bytes.

Mutate

tag

Tags are library folder membership. Comma-separated lists accepted. Both --add and --remove are optional (at least one required).

update: patch metadata

Tri-state per field: pass --set-<field> to set, --clear-<field> to null it, omit to leave alone. Fields:
  • --set-name / --clear-...
  • --set-published-at <YYYY-MM-DD>: publication date.
  • --set-effective-at <YYYY-MM-DD>: valid-from date (feeds G1 as_of retrieval).
  • --set-effective-until <YYYY-MM-DD>: valid-until (exclusive). Old contracts still match earlier as_of values; NULL / cleared = still current.
  • --set-issuer <text>: who issued the doc.
  • --set-revision <text>: version label.
The effective_at / effective_until pair is what makes time-aware retrieval work. Set them on regulated / dated documents so as_of queries land on the right version.

supersede

Marks the document as superseded: excluded from every search path but still queryable for audit. The version_misrank retrieval fix (loop-spec §F.2). --restore reverses it. Prefer supersede over delete: reversible, keeps the audit trail.

delete

Removes the document row and its chunks. S3 bytes are garbage-collected out-of-band. External-URL rows just drop the reference. Irreversible. Supersede is almost always what you want instead.

KG grafting

extract-structure

Runs the LLM structure-extractor on the document, grafting hierarchical nodes / edges into a causal graph instance. Omit --kg-instance to mint a fresh KG named after the document.

Recipes

Bulk-ingest a Dropbox export

Move a set of docs from one folder to another

Backfill regulatory dates on already-ingested policies

policy-dates.tsv has doc_id<TAB>YYYY-MM-DD per line.

Snapshot document metadata to git

Tracks library changes without exporting content.

Audit tag coverage

Prints untagged documents: candidates for auto-tagging via a Foundry enricher.

Preview the first chunks of every doc in a folder

Where document IDs come from

  • App every library document shows its ID in the details panel.
  • Foundry logs pipelines logs get shows IDs of documents a run produced.
  • documents list for scripting.