Understand
Two hubs: extraction and judgement
A transcript is raw material, not knowledge. The distinction between the backend that extracts text and the framework that decides what it means.

Most people arrive at this already believing they know what it is. They have been told about an AI that writes up meetings, and they have decided that is not interesting.
They are right that it is not interesting. They are wrong that it is what this does.
The distinction
Two things happen to a conversation, in sequence, and they are done by different parts for different reasons.
| An ingestor | core-skills (in the session) | |
|---|---|---|
| Where | a service, a local model, or your clipboard | inside a Claude Code session |
| Job | extracts from a conversation — speech to text, speakers, searchability | processes on the user’s behalf — judges, decides, writes into the vault |
| Works on | audio and raw conversation | text that already exists |
| Knows about | one conversation | the project, the people, the history, the contract |
| Produces | a transcript | meeting documents, tasks, insights, a recap |
An ingestor does the mechanical part. Audio or a meeting in, text out. It makes no decisions
about your vault and does not know that a task ledger exists. Pasting a .vtt file you
exported from Teams is an ingestor — the simplest one, and it needs no account, no
integration and no service in the middle.
core-skills does the judgement. A transcript is raw material, not knowledge. What turns
_inbox/260423-meeting.md into a document with decisions, owned tasks and a recap is a
skill running in a session, with the contract, the folder’s own configuration and earlier
insights in its context.
The test: swap one out
The difference shows most clearly in what breaks.
Swap the ingestor — a different vendor, a local Whisper, a platform’s own transcript, or nothing at all except a file you paste — and the chain still works. A better or worse transcript, the same processing afterwards. That is why the cache contract is source-agnostic, and why a meeting platform’s built-in transcription is a fallback rather than a threat.
Swap core-skills and there is no chain left. What remains is a pile of transcripts.
Which is why calling the backend “the hub” is wrong
It is a tempting shorthand and it hides the only boundary in the whole chain that matters. The ingestor is replaceable by design. Several providers expose transcripts over MCP — Deep Thought, Klang and Fireflies among them — and are interchangeable in that socket. A local model with no network is another. So is a file you drag into a folder.
The part that is not replaceable is the part that knows what a decision looks like, whose task it is, and what belongs in a recap. That part is not a service. It is a set of instructions loaded into a session on the reader’s own machine, against their own files.
What this means for a meeting recorded somewhere else
A meeting recorded in a conferencing platform and fetched down enters exactly the same chain as a voice memo from a phone: the extraction step turns it into text, and the framework turns it into something the vault can act on.
That is the whole reason such a recording is placed in the existing audio inbox rather than given a fourth track of its own. A source that behaves like the other sources can be read by something that does not know which source it is.
From ecosystem.yaml, contract 41 · core-skills 1.89.4 · read at build 2026-10-08