Kit is a private beta, and a memory you cannot check is worth nothing, so this is the list we would want if we were deciding whether to try it. Everything here is known, reproducible, and ours to fix. If you find something that is not on this list, we would like to hear it.
Live demo and evidence status checked 9 September 2026. Other product limitations retain their last documented scope.
On 9 September the public 17-question exam passed 12 questions, including 0 of 4 hard absence questions. The token-overlap baseline passed 12, with two supersession questions marked not applicable. Both missed the removal-date lookup. These totals do not make the systems equivalent: Kit passed the two graph checks and the baseline passed two more absence checks. Read the dated results and method.
The later 120-question report uses a private repository and sealed questions. Its published page explains the method but does not provide the files needed to reproduce its headline. The 17-question public exam is a separate study.
In a live check on 9 September, “When did Lady Whiskerdown land on the moon?”
returned five records while retrieval_meta.coverage.answerability
said absent. “Miss Prudence Inkpaw” also returned neighbours
in the Papers, although the guided demo's narrower query returned none.
The absence signal is useful, but a visitor reading only the hits can
still be misled. Ask for include_meta=true, inspect the
answerability and gap fields, and do not turn a similar record into an
answer to an unsupported proposition. The guided five-question run is
not evidence that arbitrary absence questions return empty lists.
Recall across a federation works, and each result names the Kit it came from. Following that citation back does not. Ids are local to each Kit and there is no source-qualified form yet, so asking for id 2 after a peer returned one gets you your own record number two, or nothing.
When two federated Kits assert different things, both come back and both are marked contested. That part is deliberate, because a memory that silently picks a winner is worse than one that shows you the argument. It is also unfinished: there is no trust-per-claim policy yet, so deciding which source to believe is still your job.
They are seeded archives with the questions, rubric and explanatory answers written by us in advance. Explanations are released after live evidence checks, not generated by a model. They show that memory can be stored, scoped, corrected and retrieved live. They do not demonstrate ingestion, nightly consolidation, decay, handoffs, backup or restore, because those happen inside a running Kit over weeks. Use a bounded trial to check those behaviours in the app. No published customer study establishes the benefit for your workflow.
Every provenance column in those archives is also empty, because the entries were written for the demonstration and no session or commit produced them. The ballast archive is the one where each entry carries a real, clickable source URL.
The two demonstration Kits have their own databases, their own keys and their own processes, and neither can read the other. They do run on one host. The boundary being demonstrated is the software one.
Claude Code, Codex, Cursor, VS Code and the other MCP agents are captured live as you work; browser chat capture is still in development. Anything that speaks MCP can talk to Kit, and being able to talk to it is not the same as having your history captured from it. The supported page marks which is which.
Source ships inside the app so you can read what runs on your machine. It is not open source, and nobody outside this project has audited it. There is no SBOM and no published threat model yet. The download is signed, notarised and checksummed, which is a different and much smaller claim.
The release notes on the install page are complete rather than curated, so they include the month Telegram was broken on every personal install, the reflection cycle that had not run since the runtime migration, and imported text that could execute code inside the app. All fixed, all recent. The rate of serious fixes is the honest signal about maturity, and it is why the advice below is what it is.
What we would tell a friend: try it on one real but non-critical project, keep the original sources, and check what it remembered rather than assuming. Do not make it the only copy of anything yet.
Kept here because a list of what is wrong is only believable beside a list of what was wrong.
A key limited to one knowledge area was correctly limited on its own Kit, and peer results were merged without that limit applied, so a federated query could return areas the key had never been granted. This is the mechanism Kit offers for sharing one area with a colleague’s Kit, so it mattered a great deal more than the fictional data it exposed.
Found by an unaffiliated AI agent evaluating the public demo. Reported, reproduced, fixed and verified the same day. Peer results now face the same scope check the local ones do, on our side of the connection, and tests pin every merge site so it cannot come back quietly.
The demonstration said one of its questions could not be answered from either archive alone. A reviewer read both and showed that one of them does most of the work, so it now says what the second archive actually adds. The same review found a placeholder shipping unsubstituted inside an answer, and an answer that stopped one sentence before the fact that completed it.