remember(), Khora runs a three-phase pipeline: it checks whether the
content is new, splits and analyzes it, and optionally connects it to what’s already
stored. This page is the conceptual tour of that write path. For exact signatures see
the API reference.
Phase 1: Staging
Before any expensive work, Khora asks “have we seen this?” It computes a SHA-256 checksum of the content and skips the document if that checksum already exists in the namespace, so re-uploading the same content (even under a new title) is a no-op. It also resolves the source timestamp (when the content originated, not when it was ingested), which powers temporal recall and becomes each chunk’soccurred_at.
Set it explicitly (recommended). Your connector knows its source, so resolve the
one meaningful instant and pass it. It’s unambiguous, and it takes precedence over the
fallback below (ISO-8601 strings are coerced for you):
source_timestamp, Khora scans the
document’s metadata for a set of recognized timestamps field and takes the first present:
sent_at → created_at → timestamp → date → occurred_at → started_at → updated_at.
This is a heuristic shortcut. Khora doesn’t know the data came from Slack or a
calendar, it only maps the field names the connector wrote, so set the one that fits
your source (sent_at for Slack/Gmail, occurred_at for calendar/events, created_at
for issues). Matching is case/separator-insensitive (occurredAt / occurred-at /
OCCURRED_AT all resolve to occurred_at, and an exact snake_case key wins on a tie).
The one special-case: setting source_type to calendar / meeting / event in
metadata flips the order to prefer event time (occurred_at) over dispatch time
(sent_at).
Custom metadata and provenance
metadata is a free-form per-document dict for anything your application needs to carry
alongside the content. Khora stores it on the document and denormalizes it onto every chunk,
so you can gate recall on it later with a recall filter,
down to a nested field by dotted path (filter={"metadata.team": "ingest"}).
The provenance kwargs you set at ingest map to the filterable system keys: source_name,
source_type, source_url, source, title, external_id, and source_timestamp.
Phase 2: Enrichment
This is where content becomes knowledge. Khora uses a staged batch architecture: every document is chunked first, then embedding and extraction run concurrently (asyncio.gather: extraction doesn’t need embeddings and vice versa), then results
are written in batches.
Chunking
Documents are split into focused, embeddable pieces. Pick a strategy with thechunk_strategy kwarg. Size and overlap are set globally via
KhoraConfig.pipelines.chunk_size / chunk_overlap (defaults 512 / 50 tokens):
Each chunk records which chunker produced it in
chunker_info.
Embedding
Each chunk is converted to a vector (defaulttext-embedding-3-small, 1536-dim) via
LiteLLM, so any provider works. Embedding is batched (up to ~200 texts per call,
sub-batches running concurrently), not one API call per chunk.
Extraction
An LLM reads each chunk and extracts entities and relationships. Two kwargs are required on everyremember(). They tell the extractor which types to look for:
ExpertiseConfig via expertise=.
See Expertise & ontologies for how to build and apply one,
or the resume-search workload for a runnable example. This way, you can restrict LLM freedom in what it extracts and treat ontologies more like a hard schema then guidance.
Selective extraction (cost control). By default Khora doesn’t send every chunk to
the LLM. On the default VectorCypher engine, skeleton PageRank ranks chunks by
keyword-graph importance and sends the top
skeleton_core_ratio fraction (default 0.50)
to full LLM extraction. The chunks it skips get no graph edges, but stay retrievable
through the vector (and, when enabled, keyword) channels. The generic ingest_documents()
path uses a different selector: a ChunkImportanceScorer scoring entity density (35%),
information density (25%), position (20%), and length (20%), which keeps the top
extraction_importance_ratio (default 0.7) and gives the rest lightweight
CO_OCCURS_WITH edges. Tune the VectorCypher selectivity via skeleton_core_ratio (see
VectorCypher).Co-occurrence edges
When the LLM extracts several entities from a chunk but doesn’t state an explicit relationship between two of them, Khora still links them with a weak co-occurrence edge, on the assumption that things mentioned together are usually related. The point is connectivity: VectorCypher answers questions by traversing relationships, so an entity with no edges is invisible to graph recall. Co-occurrence makes sure nothing is left stranded. It also captures the implicit “these were discussed together” links the model didn’t bother to name. There are two kinds, depending on where they come from:
Both are deliberately low-confidence, so they never outrank real, LLM-extracted
relationships. The dream phase down-weights them (to
0.2) and can
prune them.
Keep it on (the default) when you want forgiving recall over messy, real-world data:
better to over-connect and let related-but-unstated entities surface than to miss a link.
Turn it down when you want a graph that mirrors your ontology cleanly. On VectorCypher
the ASSOCIATED_WITH densification is always on (there is no config switch), but you can:
- prune low-confidence / co-occurrence edges in the dream phase
(
KHORA_DREAM_OPS_PRUNE_EDGES) when “edge soup” starts to hurt retrieval, and - drop the separate event edges (
EVENTentities +PARTICIPATED_IN) withVectorCypherConfig(store_events=False). See the ontology example for a worked run.
Phase 3: Expansion (optional)
After enrichment, Khora can connect the new content to the existing graph:- Entity unification: the same entity written different ways (“Microsoft Corporation”, “Microsoft”, “MSFT”) is merged via exact, fuzzy (edit-distance), and embedding matching.
- Relationship inference: new edges derived from existing ones (Alice and Bob both
WORKS_FORAcme → AliceCOLLEAGUE_OFBob), driven by the expertise config.
Three ways to ingest
Error handling
A failure in one document never fails the batch. Errors are captured per document and surfaced in the result counts:embedding_model and extraction_model are not per-call kwargs. Set them at
construction time via KhoraConfig (KHORA_LLM_EMBEDDING_MODEL, KHORA_LLM_MODEL).
The per-call kwargs are chunk_size, chunk_strategy, entity_types,
relationship_types, expertise, source_timestamp, metadata, session_id, and
external_id. Passing chunk_size=None falls back to KHORA_PIPELINES_CHUNK_SIZE.Retrieval
The read path: how
recall() finds and ranks what you ingested.Core APIs example
Runnable
remember_batch, ontology config, and entity reads.