The Anki backbone

DRAFT — will change

Every word and phrase you meet in the chat gets tracked, with a status — automatically, from conversation, never hand-authored. A read-only web page lets you see everything you've picked up and where each item stands. It's Anki's memory model with Anki's worst tax removed. This is the spine of V1 — the teaching layer that builds on V0 onboarding.

7 min
Versioning. What was first drafted as "V1" — the sticky first conversation — is now V0 (design · plan). V1 is the teaching layer: this Anki backbone — full FSRS scheduling, the tracking-and-timing spine — plus themed clusters, a reference layer that groups related words so a due item has neighbors to lean on. What's still later: skill rating, anti-annoying cadence, other languages.
  1. The thesis
  2. Where the backbone sits
  3. The backbone advises; the agent decides
  4. What gets tracked — the item
  5. The status model
  6. Capture: extraction, not authoring
  7. The web page (read-only)
  8. Why you can't add words by hand
  9. Anki, minus its taxes
  10. Where it lives
  11. V1 vs later
  12. Open questions

The thesis

Onboarding hooks them. The Anki backbone is what makes it a learning tool and not just a charming chat. Underneath the conversation runs a quiet ledger: every word and phrase the user meets, each with a status — have they just seen it, are they wrestling with it, can they use it cold? The user never fills out a card. They just talk, and the ledger fills itself. Then a web page hands it back to them: here's everything you've picked up.

This is the same bet Anki makes — catch a memory right before it fades — but with the two things that make Anki a chore deleted: you don't author cards (the conversation extracts them) and you don't self-grade (the conversation reveals the status). It tracks your own life's words, not a downloaded deck.

Where the backbone sits

V0 · Onboarding

A sticky first conversation that hooks them and gets them producing Japanese naturally — the prerequisite, not the teaching. V0 design & plan.

V1 · Teaching

The actual learning layer, in two halves: this backbone (track every word/phrase + status + FSRS schedule, auto, from the chat, with a read-only web page) and themed clusters (group related words as reference → proactive themed opens).

The backbone is the data-and-timing spine: it tracks every item, grades it from the chat, and runs FSRS over it to decide what's due. The themed-clusters doc is the reference layer on top — grouping due items into themes so the agent has related words on hand when it resurfaces one. Both are V1. What's deferred past V1 is only a skill rating and other languages.

Japanese-only, like the rest of V0/V1. No language picker; readings, statuses and the web page all assume Japanese. Other languages are post-V1.

The backbone advises; the agent decides

The backbone is a reference, not a director. It tells the agent what the user knows, what's gone shaky, and what's worth reusing — but it never grabs the wheel. The agent owns conversation flow. Its objective is to consult the ledger and work its items in where the chat gives a natural opening — never to robotically throw a tracked word in out of nowhere because something is "due."

Anti-pattern. The ledger flags 久しぶり as shaky, so the agent shoehorns it into an unrelated message. That's a flashcard popping up mid-conversation — the exact feel the whole backbone exists to kill.
The rule. Naturalness beats coverage. A due item that doesn't fit waits for its moment; the conversation never bends to serve the ledger. The user mentions not seeing a friend in ages — then 久しぶり falls in on its own.

Holds whether the agent is reading status mid-chat to pick what to reuse, or opening a themed proactive message — it surfaces an FSRS-due item only when the thread opens a door, never as a drill. FSRS decides what's due; the agent's judgment still decides when it actually lands, always.

What gets tracked — the item

The unit is an item: one word, phrase, or small grammar pattern the user has met. Kept per learner, alongside the existing profile.

FieldExampleWhy
surfaceっぽいThe form as it appears.
readingっぽい · 日本(にほん)Kana reading — so the page is always readable, same rule as the chat.
gloss"kinda / -ish / seems like"A short meaning, in their native language.
statuslearningWhere this item sits — the state machine below.
firstSeen / lastSeenJun 21 / Jun 23When it entered, when it last came up — inputs to FSRS.
encounters4How many times it's surfaced — exposure count.
fsrsdue Jun 25 · stability 4.2 · difficulty 0.3The live FSRS schedule — next-review date, memory stability & difficulty. Recomputed each time the chat grades the item.
context"talking about being tired"Where it came from — the your own life hook. Not a dictionary line.
nearMissボダリング → ボルダリングThe almost-right version, if there was one.

This is the Progress store made real — the one genuinely domain-specific memory store. Profile stays flat facts; Voice stays a small descriptor.

The status model

Four states, lifted from Anki's New → Learning → Review → Relearning loop but renamed for what they mean here, and — crucially — inferred from the conversation, not chosen by the user pressing a button.

New Learning Known Shaky re-meets uses it cold missed it keeps using
An item enters New, gets drilled by re-exposure into Learning, graduates to Known when used unprompted, and drops to Shaky if it later misfires — exactly Anki's loop, just driven by what the user actually does in the thread.
StatusMeansHow the agent infers it
newJust introducedThe agent (or the user) used it for the first time; user hasn't produced it yet.
learningWrestling with itSeen a few times; followed it with help, or produced it with a near-miss / gloss.
knownOwns itProduced it correctly, unprompted, more than once. No gloss needed.
shakySlippingWas Known, but recently missed it, needed the gloss again, or hasn't used it in a long while.

No four buttons. Anki's weak point is the honest self-grade — garbage-in self-grading corrupts every interval. Here the grade is the conversation: did they follow it, produce it, fumble it, fall back to English? The agent reads that and moves the status. The user never rates anything.

The status is the human-facing label; underneath, that same inferred read becomes an FSRS grade (again / hard / good / easy) that updates the item's stability and next-review date. So the friendly four-state badge and the real scheduler are driven by one thing — how the user actually handled the word. FSRS is exactly Anki's scheduler; we've only swapped the button-press for the conversation.

Capture: extraction, not authoring

After a turn, the agent notices vocab events and writes them to the store — the single most differentiated move, because it deletes the work Anki dumps on the user.

1
New item. Agent drops っぽい naturally → create an item, status new, with the context it came up in.
2
Promotion. User later writes 疲れてるっぽい unprompted and correct → bump toward known, bump encounters, stamp lastSeen, and feed FSRS a good grade so its next review moves out.
3
Near-miss. User says ボダリング for ボルダリング → record the near-miss, hold at learning. (Catching near-miss vs. typo is an open question, shared with the design doc.)
4
Lapse. A known item needs a gloss again, or goes long unused → an again grade drops it to shaky and FSRS pulls its next review in — the cue for a themed resurfacing.

Mechanically this is an extension of the existing remember tool (see where it lives) — a structured items write, not free text. The agent already jots facts; this gives it a typed slot for vocab.

The web page (read-only)

The user-facing surface of the backbone: one page, your whole vocabulary, grouped by status. A trophy case and a progress mirror — not an editor.

あなたの言葉 read-only
23 words · 6 known · 11 learning · 4 new · 2 shaky
All Known Learning New Shaky
っぽい-ppoi
kinda / -ish · "talking about being tired"
learning
seen 4× · 2d ago
ボルダリングborudaringu
bouldering · said it as ボダリング first
learning
seen 3× · 2d ago
大好きだいすき
love / really like · "loves Japan since a kid"
known
seen 7× · today
行き詰まるいきづまる
to be stuck · came from "I feel stuck"
new
seen 1× · today
久しぶりひさしぶり
long time no see · used it, then blanked
shaky
seen 5× · 9d ago

Every row is something they actually said or met — readings attached (so kanji is always decodable), the gloss carries the context it came from, and the status pill shows where it stands. Filter chips slice by status. That's the entire V1 page: view, don't edit.

Why you can't add words by hand

The web page is deliberately read-only — no "+ add word" in V1. This is a real design choice, not a missing feature, and it's load-bearing for the whole thesis.

Manual authoring is the tax we're deleting. Hand-building cards is Anki's single biggest source of failure — bad cards become leeches. A "+ add" button drags that exact tax back in. The backbone's whole pitch is that you never author; let the form in and we've reinvented Anki.
A typed-in word is a hollow item. Vocab earns its place by appearing in real talk — that's what gives it a context line, an encounter count, a near-miss, a status the agent can trust. A word punched into a form has none of that. It's a card, not a memory of your own life.
You already have an input: talking. Want a word tracked? Use it in the chat. The conversation is the add button — and it's a better one, because it also captures how you used it.

Noted as "for now": a later milestone may let you nudge an item (mark something known, or ask the agent to start working a word in). But it routes through the conversation/agent, never a raw deck-builder form. The door's open; the form stays shut.

Anki, minus its taxes

The backbone keeps Anki's memory machinery and drops everything Anki offloads onto the user. Side by side:

AnkiThe Anki backbone
You author notes — fields, templates, the Twenty RulesThe agent extracts items from conversation — zero authoring
You self-grade with four buttons each reviewStatus inferred from how you actually used the word
Cards are downloaded decks / abstract factsItems are your own words, with the context they came from
Decks & queues are the retrieval UIStatus groups on one read-only page
You go to the app to reviewReview comes to you in the thread (themed proactive sessions — V1)
FSRS scheduler picks the next review dateThe same FSRS, in V1 — fed by inferred grades instead of button presses

The scheduler is the easy, borrowable part; FSRS is open-source and battle-tested. The differentiated work is everything above it — capture and grading without the user lifting a finger.

Where it lives

Today memory is a flat free-text notes blob in core/.memory.json, keyed by handle. The backbone adds a structured store next to it:

learner = {
  profile: { … },              // unchanged: name, level, scripts, goal…
  items: {                     // NEW — the Anki backbone
    "っぽい": {
      surface: "っぽい", reading: "っぽい", gloss: "kinda / -ish",
      status: "learning", encounters: 4,
      firstSeen: "2026-06-21", lastSeen: "2026-06-23",
      fsrs: { due: "2026-06-25", stability: 4.2, difficulty: 0.3 },
      context: "talking about being tired",
      nearMiss: null
    },
    …
  },
  history: [ … ], summary: "…"  // unchanged
}

The agent writes items through an extended remember tool (a typed vocab slot beside the existing note/profile fields). The web page reads the same JSON. Single-user dogfood scale — a JSON file is fine; swap for SQLite/Supabase if it grows, same as the rest of memory. This is the Progress store — the load-bearing memory store.

V1 vs later

V1 — now (this doc)

  • Structured items store per learner.
  • Automatic capture from the chat (extended remember).
  • Inferred status — new / learning / known / shaky.
  • Full FSRS scheduling — inferred grades drive per-item timing & decay.
  • A read-only web page to view all vocab by status.
  • Feeds the themed clusters reference layer — also V1.

Later (post-V1)

  • Skill rating (ELO) to drive dosage smoothly.
  • Anti-annoying cadence tuned by data, not feel.
  • Any manual nudge — and only via the agent, never a deck form.
  • Other languages beyond Japanese.

Open questions