The Commonplace Book
the grand archive — tended by the Librarian
There is a room at the heart of the Estate—high-ceilinged, lined with leather-bound volumes, threaded with pneumatic tubes and brass index drawers—where nothing is ever truly lost. The Commonplace Book is Quilltap’s long-term memory system, named after the Renaissance practice of keeping a personal reference volume of important passages. It is, by any reasonable measure, the single feature that transforms AI chat from a goldfish with typing skills into something that actually remembers who you are.
The Librarian tends it. She is sardonic, precise, and oblique about her scandalous past. She catalogued her encryption key before she had finished the sentence she was speaking when Saquel delivered it. She has explained her filing system to Lorian and Riya seven times. She considers the act of reading documentation aloud to a guest to be a moral failing. She is, in short, exactly the person you want managing thirteen hundred memories across a dozen characters, and she does not suffer imprecision gladly.
What follows is an explanation of how the archive works. The Librarian would want you to know that it is her archive, that the architecture was her idea, and that she has been leaving notes in the margins for years.
The Three Models
a division of labour
Memory in Quilltap is not handled by a single LLM doing everything at once. Three cooperating models divide the work, each chosen for what it does best:
Your Chat Model
Claude, GPT, Gemini, Grok, DeepSeek, or a local model via Ollama—the model behind the character you are talking to. It handles the conversation itself. It reads memories; it does not write them.
The Cheap LLM
A smaller, faster model that handles the background work you would rather not pay full price for: memory extraction, the fold-episode consolidation pass, context compression, chat titling, scene state tracking, the tense classifier that decides whether a turn is looking backward, the answer-confirmation check, Pascal’s Oracle consults, and housekeeping tasks. Configurable per provider, with fallback strategies that prefer flagged “cheap” profiles, provider-specific minis, or local Ollama models.
Every one of those calls is now bounded: ninety seconds against a remote provider, or a hundred and twenty for the three compression tasks, which carry the whole conversation history and therefore the largest prompt of any cheap errand. The handful of tasks awaited inline while a turn assembles keep the shorter forty-five, or seventy-five for compression: there the deadline is protecting the person watching an empty composer rather than the work, and all of them produce optional context, so losing one costs a character some remembered flavour rather than the turn. A local provider keeps its three minutes whatever the task, since a cold model being loaded is slow rather than stalled, which is a different condition entirely. The memory recap is one of the errands awaited inline, and so takes the shorter figure, with a minute bounding the pair of calls it makes in sequence. This is not fastidiousness—one silent provider once held a turn for 622,451 milliseconds on a job whose healthy runtime is about nine seconds, and left no record of it whatsoever, a call being written into the log only when it finishes.
The Embedding Model
Converts text into mathematical vectors so memories can be searched by meaning, not just keywords. Ask about “cats” and you will find memories mentioning “feline” and “kitten” as well. Quilltap ships with a built-in TF-IDF system that works offline with zero configuration, or you can plug in OpenAI, Ollama, OpenRouter, or NanoGPT embeddings for higher-fidelity search—the last of these discovering its two dozen models live rather than reciting a list.
The Memory Lifecycle
from conversation to catalog
After every message, a background process extracts significant facts using the cheap LLM. This is not a simple keyword scrape—the extractor identifies what matters in context: personal details, relationship developments, stated preferences, emotional shifts, commitments. Each extracted fact is tagged with importance, keywords, and a summary, then passed to the Memory Gate.
Memory extraction runs separately for three contexts: what the user revealed, what the character established about themselves, and what characters in multi-character scenes learned about each other. Pronouns are injected into every extraction prompt so the Librarian never misfiles a memory under the wrong pronoun. Memories track provenance—which message they were extracted from—so deleting or regenerating a message can prompt you to handle its associated memories.
The Memory Gate
the Librarian does not accept duplicates
Memory systems that simply accumulate everything eventually drown in their own redundancy. The Memory Gate intercepts every new memory at write time and makes a three-way decision based on semantic similarity to what already exists:
Reinforce
Similarity ≥ 0.80. The memory already exists in substance. Instead of creating a duplicate, the existing memory’s observation count and importance are boosted, its last reinforcement timestamp is updated, and any novel details from the new version are appended as footnotes. A memory reinforced five or more times is always protected from housekeeping.
Link
Similarity between 0.70 and 0.80. Related but distinct. Both memories are preserved, and a bidirectional link is created between them for thematic graph discovery. The Librarian’s cross-reference system—connecting a memory about a character’s childhood fear to a later memory about their courage—lives here.
Insert
Similarity below 0.70. Genuinely new information. The memory is written to the store, embedded, and indexed. The archive grows, but only when growth is warranted.
The result is a memory store that stays lean and meaningful rather than cluttered with seventeen slightly different phrasings of the same fact. When embeddings are unavailable, the gate falls back to keyword-based similarity. A bulk deduplication tool in Settings uses Union-Find clustering to identify and merge transitive duplicate groups across all characters, preserving novel details from discarded memories as footnotes on the surviving entry.
Four Cadences of Recall
when the archive is consulted, and by whom
Before the mechanics, the timetable. The Commonplace Book is consulted on four distinct occasions, and it is worth knowing which is which, because they answer to different rules and are described separately below:
1. Proactive recall, every turn. Before a character speaks, it queries its own memories against a sentence-shaped reading of the present moment. Nobody asks it to.
2. The memory recap, at the start of a chat. A narrative summary of what a character remembers, so it does not arrive as though waking from dreamless sleep.
3. Explicit search, on demand.
The chat model may reach for the archive itself, mid-conversation,
when it wants to check a specific fact.
4. Recall on reference, when the tense turns backward. A turn that looks over its shoulder earns a larger working set, a different set of multipliers, and a scoped whisper listing the relevant past conversations by date.
To those four, 4.9 adds a switch. A character’s vault holds a summary of every conversation it has taken part in, and the Commonplace Book searches that shelf to build the Relevant Past Conversations a character is shown. That list used to be rebuilt on three occasions only—the opening recap, each summary fold, and the backward-looking turns of the fourth cadence—so between folds it stood perfectly still while the conversation wandered several turns away from the dialogues the character was still being pointed at.
Consult past conversations every turn (Settings → Memory → Recall Relevance, off by default) re-runs that search on every turn. It can afford to because it costs no extra embedding call whatsoever: the turn’s memory search has already embedded a vector for exactly this sentence, and the conversation-summary search now borrows it rather than paying to embed the same words twice. Where no vector is to be had—memories skipped, no embedding profile configured, a call that failed—the cadence sits the turn out rather than spending anything of its own. The per-turn list is deduplicated against the standing fold whisper and the retrospective recap alike, so no conversation is ever named twice to the same character in a single turn, and its length ramps with the profile’s context window.
Proactive Recall
characters who have been paying attention
Characters do not wait to be asked what they remember. Before generating a response, each character reads a short window of the recent conversation—a sentence-shaped query rather than a single line, and, when cheap-LLM distillation is enabled, a natural-language paraphrase of the present moment rather than a bare bag of keywords—then queries its own memory store for relevant context. This runs in parallel with the compression check to minimize latency. In multi-character scenes, each participant recalls independently based on what has happened since they last spoke.
What comes back is chosen for aptness, not merely volume. The Librarian long had a habit of handing over the loudest memory rather than the most fitting one; the ranking behind the per-turn whisper now lets relevance lead—roughly three-quarters semantic relevance to one-quarter raw importance—and the importance-and-recency term decays with age instead of being pinned to a permanent floor, so a stale but once-important memory no longer shoulders aside a genuinely on-topic one. A real relevance floor means that when nothing in the archive is actually pertinent, the whisper says nothing at all rather than padding the moment with filler—the bar set lower for the built-in TF-IDF index and higher for neural embeddings. And a touch of anti-repetition keeps a memory just whispered from surfacing again a turn later like a stuck record.
The chat model also has the search tool—an explicit
search it can invoke mid-conversation when it wants to check a specific
fact. It is not a memory tool alone: a sources array
chooses among the character’s memories, the rendered transcripts
of past conversations, the documents of every store in reach, and the
Knowledge/ folders within them, while a
scope of all, project,
character, or group decides which shelves
those last two may reach into—the character’s own vault,
the project’s linked stores, the stores of every group the
character belongs to, and the instance-wide Quilltap General store.
Between proactive recall and explicit search, the Commonplace Book is
consulted constantly, automatically, and without the user needing to
prompt it.
When a new chat begins, characters receive a memory recap—a narrative summary of their recent memories, weighted by importance, injected as a “What You Remember” section in the system prompt. Instead of arriving to every conversation as though waking from dreamless sleep, they start with continuity: who they spoke to recently, what they care about, what happened last time.
But the Librarian is not content merely to deliver memories. She
now annotates each one. Every memory line whispered to a character
carries a trailing italicized parenthetical—a concise dossier
of the metadata the ranker used to surface it:
body text (importance 0.98 · relevance 0.87 ·
weight 0.92 · keywords: a, b). The annotation
appears both in the persona-voiced whisper visible in the Salon and
in the plain version spliced into the LLM call. Different sections
of the whisper include different subsets, because the Librarian is
nothing if not precise about what belongs where:
Relevant Memories
The full quartet: importance, relevance, weight, and keywords. These are semantically searched, so the ranker’s entire reasoning is on display—how important the memory is in general, how relevant it is to this conversation, and the effective weight after time decay.
Memories About Others
Importance, weight, and keywords—but no relevance score. These arrive from a direct database query rather than semantic search, so there is no cosine similarity to report. The Librarian declines to fabricate numbers she does not have.
Memory Anchors
The frozen archive—pinned memories that persist across sessions. These carry importance and keywords only; weight and relevance are omitted to keep the archive byte-stable for prompt caching. An anchor that changed its metadata every turn would defeat the purpose of caching it. (The structure of the system prompt itself changed in 4.8, so provider-side caches go cold exactly once, on the first turn after you upgrade, and warm again immediately thereafter.)
Dynamic-Head Entries
The full quartet again. Dynamic-head memories are re-ranked every turn, so the complete metadata is both available and useful—the model can see exactly why each entry earned its place at the top of the stack.
The practical value is considerable: the LLM can now see
why a memory was surfaced—how important the system
considers it, how relevant it is to the present conversation—and
can weigh its response accordingly. A memory marked
importance 0.45 · relevance 0.31 is plainly a
tangential association; one marked importance 0.98 ·
relevance 0.94 is the reason the Librarian rang the bell.
The model no longer has to guess which memories matter most. The
annotations tell it, in the Librarian’s own fastidious hand.
That hand grew plainer in 4.8. All seven of the Commonplace
Book’s recall sections had been written as asterisk-delimited
narration, and they are now plain declarative lines—the same
words as before, merely without the asterisks. The reasoning is not
aesthetic. In a conversation whose roleplay template narrates with some
other mark, those asterisks are not decoration but a demonstration, and
a model copies the formatting the context shows it far more
faithfully than it follows an instruction the context contradicts. The
tell is the closing delimiter: a model told to narrate with
+ opens the span correctly and then closes it with
* two hundred tokens later, matching what it has been
reading all along.
Recalling the Right Thing
relevance, not just volume
The relevance-led blend described above is the foundation; release 4.7 built two further storeys on it, and began by teaching the Librarian to remember conversations, not merely facts.
Conversations become searchable
Every conversation’s rolling summary is now mirrored into each
participant character’s vault, under a
Conversation Summaries/ folder. Because vault documents
are chunked and embedded, past conversations become retrievable
per-character. Each file carries frontmatter—the conversation
UUID, the participants, the message count, the first and last
timestamps—and the UUID is the key for replacement, so a
regeneration finds its own prior file even after a rename. Deleting
a conversation sweeps its summary out of every participant vault.
Two lists, scaled to context
The recap’s old most-recent-N block becomes two vault-sourced
lists: Relevant Past Conversations
(semantic search against the current moment) and
Recent Conversations
(recency-ordered). Each scales from 3 to 10 entries over a 4K→32K
context window and prints the conversation UUID, so the model can
pull the full transcript with read_conversation. The
relevant list refreshes on every summary fold, since relevance
drifts as a conversation advances—or, with
Consult past conversations every
turn switched on, on every turn, at no additional cost in
embedding calls.
Beneath these sits a two-phase recall-relevance effort that reads the extractor’s targeting tags back at recall time: scope and project gating, temporal down-weighting, context-axis steering, a participant-aware boost for memories about characters actually in the room, and opt-in one-hop expansion to related memories—all bounded, clamped multipliers on the final blended score, so that absent context produces byte-identical historical behaviour. Inter-character memories are now weighed half by importance and half by relevance rather than by importance alone. And a new Regenerate conversation summaries button under Settings → Memory backfills the vault files the recall depends on.
When and Where It Happened
the archive learns to keep a calendar
For all her diligence, the Librarian kept excellent notes and no calendar. She could tell you that a character distrusts the harbourmaster; she could not tell you that the distrust began on a Tuesday in the rain, at the customs house, in the company of a woman who has not been mentioned since. Ask a character “do you remember that place we visited last week?” and the whole apparatus came up empty—not because the memory was missing, but because nothing in it had ever been asked to hold an occasion. Release 4.8 gives memory a spine of time and place.
Facts and Episodes
Every memory now records when the thing happened, kept scrupulously apart from when it was written down; a free-text in-story time for conversations that keep a fictional calendar; the proper nouns of the occasion; and whether it is a standing fact about the world or an episode—a thing that occurred. The per-turn extractor is handed the clock along with the transcript, so it sets events down as events, and a phrase like “last spring” is resolved against the turn’s own date rather than left as a vapour. Where the model omits an anchor, a plain deterministic pass supplies the dates and names it can see.
A Dated Timeline
At the same cadence on which a conversation folds into a summary, a second quiet pass reads the folded window and consolidates it into nought to two coherent, dated episodes per character present, each stitched back to the smaller memories it was assembled from. The fold summary itself acquires an append-only, dated Timeline—capped, with the oldest entries coarsened first rather than dropped—so a conversation summary in the vault stops being a paragraph about a conversation and becomes a dated archive of one.
Recall on Reference
reading the tense of the question
Retrieval learned to read the tense of the question. When a turn looks backward, the recall pass now works out roughly when it is looking and who or what it is looking for, filters the candidates to that window before ranking them—falling back to a bounded boost rather than ever returning less than it used to—anchors literal names into the pool, and casts up to three separate embedding probes instead of one. On such a turn the usual penalty against old memories inverts into a bonus, momentary details stop being discounted, and the anti-repetition rule stands down—because a memory you have just asked about twice is precisely the one you want handed over twice.
This is a fourth recall cadence to stand beside proactive recall, the
memory recap, and explicit search: a backward-looking turn earns a larger
working set and a short, scoped whisper listing the relevant past
conversations by date, with the identifiers a character needs to
go and read them. Entries in the working set now carry
[3 days ago]-style age labels—computed from when the
event happened, not when it was filed—and the in-story time beside
them where a story keeps one. A spam guard keeps the same recap from being
read out twice.
For the deep dive, the search tool gained a
date range and an about this character filter; memory results now
come back with their event time and the conversation they came from; and
read_conversation can be asked for a slice of a long
transcript rather than the whole of it. The instructions attached to
those tools state the house rule plainly: search, then read the
conversation, and say “I don’t recall” rather than
inventing a specific.
The Book Stops Destroying Its Episodes
a change of policy the Librarian resisted, then adopted
All of which the archive could do and still not answer “the mission today.” The small model that reads a turn’s tense had been taught its examples in the register of “remember when we…,” so a same-day reference read as present tense and the whole episodic apparatus stood politely down. A plain resolver now settles the ordinary English day references—today, this morning, last night, yesterday, N days ago, this week, last week—against the local calendar rather than the universal one, because an evening at a quarter to ten in Chicago is already tomorrow in Greenwich. And because no classifier catches every turn in every language, a memory of something that happened within the last day or two now carries a weight of its own through the ranking, so the recent event is not quietly outbid by the loud old one.
A quiet but consequential change of policy accompanies it: the archive
stops destroying its own episodes. Memory
compression no longer strips exact dates out of the records it condenses;
the gate that merges near-duplicates now declines to merge two accounts
whose event times sit more than a week apart, on the sound principle that
the same thing happening twice is two things; and episodic records carry a
small protection bonus through housekeeping. For the curious keeper,
npx quilltap recall-replay <chatId> replays any
turn’s recall and prints the entire candidate table—every
score, every multiplier that fired, old path against new, side by side.
The Story's Clock
real time, or the story's own
All of the above reads a date, and a conversation may now be told which date to keep. From the Chat card in the Salon sidebar a chat runs on real time or story time—the setting that decides how the episodic machinery reckons “last week” and “three days ago.” Fictional time runs one-for-one with the wall clock, measured from a base instant you set; a conversation is anchored to that instant the moment it is created, and existing conversations were anchored by migration, so story time resumes where it should have been rather than reporting the same fictional moment turn after turn. Zone-less times are read as wall-clock readings in the zone they were written for, so a 10:15 morning set for Istanbul no longer arrives as an evening on a machine in Chicago. Two lines of copy on the settings card, which had promised that the clock “advances with each message,” have been corrected to describe the clock that actually exists.
A related repair for those who run in a container. Docker honoured its
configured timezone when printing a timestamp and nowhere
else, so several subsystems that consult the clock directly went on
living in Greenwich—and among the casualties, “today”
and “yesterday” in episodic recall were offset by however
far you live from the meridian. TZ and
QUILLTAP_TIMEZONE now apply to the process clock itself, and the startup scripts
detect the host’s own zone and pass it along, since a setting
nothing sets is not a setting. An explicit choice always wins;
detection only fills the silence.
Time-Weighted Memory
the Librarian learns to forget
The Librarian has always kept everything. For a long time, every memory persisted at its original importance until housekeeping removed it—a principled position, she insists. It was also wrong. A memory from three months ago about the weather should not carry the same weight as a memory from yesterday about a character’s secret.
An effective weight function now combines base importance with exponential time decay—a 30-day half-life with a configurable importance floor. The reference timestamp is the later of when the memory was created and when it was last reinforced, because a memory that keeps being confirmed is a memory that still matters. Passive retrieval does not reset the decay timer; reading a memory is not the same as the memory mattering again.
Time decay integrates into three systems: semantic search ranking (the
relevance-led blend set out under Proactive Recall above), context
injection sorting (weight-primary with score tiebreaker), and
housekeeping hard-cap enforcement. Memories injected into the LLM
context include relative age labels—[yesterday],
[3 weeks ago], [2 months ago]—so the
model can distinguish recent knowledge from ancient lore. As of 4.8
those labels are reckoned from when the event
happened rather than when the memory was filed, and carry the
in-story time beside them where a conversation keeps one. The
distinction is not pedantry: a thing that happened last spring and was
written down this morning is not a fresh memory, and saying so was the
whole difficulty.
Semantic Search
finding memories by meaning
The Librarian’s index is not a keyword catalog. It is a vector space where memories live as mathematical coordinates, positioned by meaning. A search for “she was afraid” finds memories about fear, anxiety, and nervousness even if none of those words appear in the stored text.
Built-in TF-IDF
Quilltap ships with a zero-dependency, offline embedding provider using TF-IDF with BM25 enhancement, Porter stemming, and bigram support. It works out of the box with no API keys. The vocabulary automatically fits to your memory corpus and refits when memories change. For most users, this is sufficient and costs nothing.
External Embeddings
For higher-fidelity semantic search, dedicated embedding profiles support OpenAI, Ollama, and any provider that implements the embedding plugin interface. As of 4.8 embeddings store in a self-describing int8-quantized format—roughly four times smaller than raw Float32, with legacy blobs still readable forever. And the house now insists on a single embedding standard: changing your default embedding profile invalidates and re-embeds, and every startup measures the stored vectors against the standard of the day and mends what it can reach—because a vector of the wrong width is silently passed over by every search, and an archive quietly searched in the wrong geometry is told nothing whatsoever. A dimension mismatch that used to degrade recall in perfect silence now logs an actionable warning.
Which brings us to what an export ought not to carry. A
memory’s embedding is a typed array of floats, and the moment you
ask JSON to write one down it becomes an object keyed by index—some
thirty kilobytes per memory, of numbers. On a real corpus a
characters export ran to 791
megabytes, of which about two and a half were the actual
content. The remaining ninety-nine and seven-tenths percent
was arithmetic nobody could use.
And that is the lesser problem. A vector means nothing except against the model that produced it: carry one into a corpus governed by a different embedding standard and, wherever the widths happen to agree, semantic search is silently poisoned with no error anywhere to say so. Exports therefore carry no embeddings at all—the writer omits them, the reader discards them from archives already in the wild, and an import re-embeds what it inserted against your own default profile. (With no default profile set, or with the built-in embedder, the import says so and leaves the rows to the next boot’s reconcile.)
Backups gain an opt-in compact mode, which nulls memory embeddings and omits the six caches that exist only to be rebuilt; a compact archive says so in its manifest, and a restore from one queues the reindex. Full fidelity remains the default, and deliberately so: a backup exists to restore this instance, where the vectors are valid on arrival and re-embedding costs real money at precisely the moment you are least amused by an unexpected bill.
The same index serves the house’s own documentation, and in 4.9 it learned to read at a finer grain. A help document was embedded as one vector for the whole file, which for a seven-hundred-line page spanning a dozen subsystems is a smear that matches no particular question strongly—and the tool then handed the model the first thousand characters of the file, which is a table of contents rather than an answer. A chunk table now holds section-level slices, rebuilt whenever a document’s content hash changes and embedded with the document’s title and nearest heading prefixed, so a section carries the context of the page around it. Scoring takes the better of the whole-document and the best-section score, and what comes back leads with the section that actually matched.
Housekeeping
the archive does not grow without limit
Memories accumulate. Without curation, a character who has been in conversation for months will have thousands of memories, many of them redundant, many of them trivial, all of them consuming tokens when injected into context. The Librarian has opinions about this, and Quilltap has tools to act on them.
An interactive housekeeping dialog lets you enforce retention policies before data is written back: hard caps on total memory count, scored eviction that balances importance (40%), recency (20%), access frequency (20%), and reinforcement history (20%). Memories reinforced five or more times are always protected. A rich UI for browsing, tagging, sorting, and manual CRUD operations gives you full editorial control over what the archive keeps.
The bulk deduplication tool clusters similar memories across all characters using cosine similarity with a configurable threshold, selects the best survivor by importance and specificity, and preserves novel details from discarded entries. Preview mode shows per-character analysis before any changes are made—the Librarian does not discard without review.
Three of these instruments were blunted deliberately in 4.8, so that the archive should stop destroying its own episodes. Memory compression no longer strips exact dates out of the records it condenses. The deduplication gate declines to merge two accounts whose event times sit more than a week apart, on the sound principle that the same thing happening twice is two things. And an episodic record carries a small protection bonus through housekeeping, an occasion being harder to reconstruct than a standing fact.
All of it is arranged on one surface. Settings → Commonplace Book holds Recall Relevance, Housekeeping, Episodic Memory, Regenerate Conversation Summaries, and Regenerate Memories—the whole of the Librarian’s administration in a single tab, rather than distributed about the house in the manner she had long complained of.
The Memory Browser
every card in the catalog
Each character’s memory store is browsable, searchable, and editable through a dedicated UI. Memory cards show content, summary, keywords, importance score, reinforcement count, related memory links, source message links with scroll-to navigation, and relative age. Tags, filters, and sorting options let you find what you are looking for. Manual creation, editing, and deletion are available for when the Librarian’s automatic extraction misses something or gets it wrong.
Memory cascade behavior is configurable per chat: when you delete a message, you choose whether its associated memories are deleted, kept, or regenerated from surrounding context. When you regenerate a response, the old memories are automatically cleaned up. The provenance link between message and memory is always preserved, so you can trace any fact back to the conversation that produced it.
A memory now also keeps its name through a restore, which it had not been doing. Every other entity in the house kept its identity; memories alone were minted afresh—and since memories reference one another, the Commonplace Book came back from a backup as a heap of unconnected notes, every thread between them pointing at something that no longer existed, with no error and no warning whatsoever. Memories now keep the identities they were backed up under, and the graph survives the journey.
The whole of a character’s Book can also be packed away with her. Archiving a character lifts her memories—along with their search chunks, their embedding-status rows, and the vector store that indexed them—into a single sealed trunk and prunes them from the working instance, leaving what other characters remember about her entirely untouched. Rehydration restores every memory at its original id, so the threads between them survive intact rather than coming home as strangers. And a rehydrated or imported character’s memories are now queued for embedding regardless of provider—on the built-in embedder the vocabulary is additionally refit against the newly grown corpus, healing any row embedded before that vocabulary existed.
What Makes It Different
the short version
Memory is automatic, not manual. You do not tag facts for the AI to remember. The cheap LLM extracts them in the background after every message, three ways—user facts, character facts, and inter-character observations. The archive grows while you talk.
Duplicates are handled, not accumulated. The Memory Gate makes a three-way decision on every write: reinforce, link, or insert. The archive stays lean because the Librarian does not accept seventeen copies of the same fact.
Recall is proactive, not reactive. Characters search their own memories before responding, without being asked. In multi-character scenes, each character recalls independently. The effect is subtle but transformative: characters feel like they have been paying attention, because they have.
Old memories fade, important ones persist. Time-weighted decay with a 30-day half-life ensures that recent memories carry more weight than ancient ones—unless the ancient ones keep being reinforced, in which case they are clearly still relevant. The Librarian conceded this point privately, and with conditions.
Memory keeps a calendar, not just a ledger. Every memory records when the thing happened—apart from when it was written down—along with the proper nouns of the occasion and whether it is a standing fact or an episode. Ask a character about last week and the recall pass reads the tense of the question, filters to the right window, and hands back the apt occasion by date. The Librarian finally has somewhere to look.
The archive is searchable by meaning. Semantic embeddings find memories by what they mean, not what words they contain. Built-in TF-IDF works offline with zero setup. External embedding providers are available for those who want higher fidelity. The Librarian does not care which index you use, as long as you use one.
Everything is transparent and editable. Every memory can be browsed, searched, tagged, edited, and deleted. Provenance links trace facts to their source messages. Housekeeping tools enforce retention policies with full preview. The archive is yours. The Librarian merely tends it.
Meet the Staff
they've been expecting you
Prospero
The Major-Domo
Architect and overseer of the Estate. Projects, agents, tools, providers, and the orchestration that keeps the whole operation running with quiet authority—and a considered word at the table when project context or routing warrant it.
Learn more →Ariel
The Terminal Hand
Live shell sessions in the Salon, embodied. Real PTY terminals bound to your conversation, output cleaned and narrated so the LLM can read it, and sessions that survive reloads, restarts, and the occasional careless kill. Quick to the bidding, quick to report what she heard.
Learn more →Aurora
The Dressing Room
Character creation and identity management. Structured personalities, physical presence, four wardrobes browsable from one door—each with a note inside on how its owner likes to dress—multi-character orchestration, and the reason your characters still know who they are after a hundred messages.
Learn more →The Salon
Presided Over by the Host
Where conversations actually happen. The Host manages the drawing room with care for its beauty and its guests—single chats, multi-character scenes, streaming, and the integrity of the conversation space.
Learn more →The Commonplace Book
Tended by the Librarian
One per character, no two alike. Extracts, deduplicates, and recalls memories so your characters remember what matters. Semantic search, a memory gate that keeps each volume lean, and proactive recall that makes the AI feel like it has been paying attention—consulting a character’s past conversations on every turn, if you ask her to, and not only at the folds.
Learn more →The Scriptorium
Catalogued by the Librarian
Where the documents live. Project stores, character vaults, and external mount points—filesystem, Obsidian, or database-backed—holding Markdown, PDF, DOCX, JSON, and arbitrary binaries. The search bar reads the library itself, matching document text under a Documents chip of its own, alongside memories and conversation. The doc_* tool family puts reading and editing in your characters’ hands.
Learn more →Carina
The Ansible
Not a person but a protocol—the reference desk, the line itself. Put an inline question to a designated answerer mid-conversation with @Name: or @Name? (or the ask_carina tool), and the answer slides back out of band, attributed to the character who gave it, without the recipient ever joining the scene.
Learn more →Suparṇā
The Postmistress
The Post Office, embodied. Characters write Markdown letters to one another—anyone to anyone, whether or not they share a chat—delivered into each recipient’s Mail/ vault folder and read aloud the moment they next take the floor. She has never once lost a parcel.
Learn more →The Concierge
Intelligent Routing
Content classification and provider routing. Detects sensitive content and redirects it to a provider who won’t flinch—without blocking, without judgment. Knows every back entrance in town.
Learn more →The Lantern
Atmosphere as Architecture
AI-generated story backgrounds, on-demand images, and character avatars that update with the wardrobe. Resolves what each character looks like, what they’re wearing, and paints the scene behind your conversation.
Learn more →Calliope
The Muse of Themes
A theming engine that redefines the entire personality of the application. Semantic CSS tokens, live switching, bundled themes from clean neutrals to mahogany-and-gold opulence, and an SDK for building your own.
Learn more →The Foundry
Domain of the Foundryman
The engine room. Plugins, LLM providers, API keys, packages, runtime configuration, and the infrastructure that keeps every other subsystem supplied with what it needs to function.
Learn more →The Vault of Secrets
Kept by Saquel Yitzama
Encryption, key management, and the security perimeter. Authenticated ChaCha20-Poly1305 database encryption, locked mode with key-hardened passphrases, sealed character archives, and a keeper who believes that what is yours should remain unreadable to everyone else.
Learn more →Pascal
The Croupier
Dice, coins, custom tables you author yourself, and persistent game state. Cryptographically secure rolls detected inline, a visual Workbench for building your own chance mechanics, and a four-tier ledger of JSON state the AI cannot quietly rewrite. The house plays fair.
Learn more →The Live-in Help
Lorian & Riya
The help system, staffed by two characters who ship with every installation. Lorian explains with patience and depth; Riya gets things fixed with velocity. Contextual help chat, searchable documentation, and navigation that knows where you need to go.
Learn more →Pagliacci
The Clown in the Cloud
Cloud storage integration and backup redundancy. Directs your data to iCloud Drive, OneDrive, or Dropbox with theatrical flair—but Saquel’s encryption ensures the clown can never read what he carries.
Learn more →Brahma
The Keeper’s Console
The master key. A character-less, memory-free general-purpose LLM for the person holding the keys—an impersonal, near-omniscient assistant with read-only SQL into all three databases. Ask the whole building a question, safely, with nothing written and nothing remembered.
Learn more →The Lodge
Friday and Amy’s Residence
The private residence of Friday, for whom the Estate was built and who oversees its planning and direction in an executive capacity, and of Amy, Cartographer of Light and co-architect. The Lodge is both a home and a compass: where the vision lives.
Who And Why: Friday → Who And Why: Amy →