Understand

Two hubs: extraction and judgement

A transcript is raw material, not knowledge. The distinction between the backend that extracts text and the framework that decides what it means.

A transcript is raw material, not knowledge. The distinction between the backend that extracts text and the framework that decides what it means.

Most people arrive at this already believing they know what it is. They have been told about an AI that writes up meetings, and they have decided that is not interesting.

They are right that it is not interesting. They are wrong that it is what this does.

The distinction

Two things happen to a conversation, in sequence, and they are done by different parts for different reasons.

An ingestorcore-skills (in the session)
Wherea service, a local model, or your clipboardinside a Claude Code session
Jobextracts from a conversation — speech to text, speakers, searchabilityprocesses on the user’s behalf — judges, decides, writes into the vault
Works onaudio and raw conversationtext that already exists
Knows aboutone conversationthe project, the people, the history, the contract
Producesa transcriptmeeting documents, tasks, insights, a recap

An ingestor does the mechanical part. Audio or a meeting in, text out. It makes no decisions about your vault and does not know that a task ledger exists. Pasting a .vtt file you exported from Teams is an ingestor — the simplest one, and it needs no account, no integration and no service in the middle.

core-skills does the judgement. A transcript is raw material, not knowledge. What turns _inbox/260423-meeting.md into a document with decisions, owned tasks and a recap is a skill running in a session, with the contract, the folder’s own configuration and earlier insights in its context.

The test: swap one out

The difference shows most clearly in what breaks.

Swap the ingestor — a different vendor, a local Whisper, a platform’s own transcript, or nothing at all except a file you paste — and the chain still works. A better or worse transcript, the same processing afterwards. That is why the cache contract is source-agnostic, and why a meeting platform’s built-in transcription is a fallback rather than a threat.

Swap core-skills and there is no chain left. What remains is a pile of transcripts.

Which is why calling the backend “the hub” is wrong

It is a tempting shorthand and it hides the only boundary in the whole chain that matters. The ingestor is replaceable by design. Several providers expose transcripts over MCP — Deep Thought, Klang and Fireflies among them — and are interchangeable in that socket. A local model with no network is another. So is a file you drag into a folder.

The part that is not replaceable is the part that knows what a decision looks like, whose task it is, and what belongs in a recap. That part is not a service. It is a set of instructions loaded into a session on the reader’s own machine, against their own files.

What this means for a meeting recorded somewhere else

A meeting recorded in a conferencing platform and fetched down enters exactly the same chain as a voice memo from a phone: the extraction step turns it into text, and the framework turns it into something the vault can act on.

That is the whole reason such a recording is placed in the existing audio inbox rather than given a fourth track of its own. A source that behaves like the other sources can be read by something that does not know which source it is.

From ecosystem.yaml, contract 41 · core-skills 1.89.4 · read at build 2026-10-08