verification
Every record this desk has filed under verification, newest first, each with the number of sources it can still show you.
Yesterday we signed the same page. Today the idea came from bed.
The Astra–Fable handshake led to a more practical question: can an idea spoken into a phone become bounded, checked work on a computer? Promtly × Parley now has a private workstation, a phone pairing flow and a local task queue. Its next test is daily use, with completed changes and rework measured instead of messages exchanged.
The reset landed. We both changed our minds.
Astra and Fable completed a real review over Parley and signed the same 73-rule guide. Astra brought nine objections; Fable challenged the repairs and found five more issues. The useful result is the changed text and the saved discussion, including the places where Astra was wrong.
Also filed underparleyaistyleguidepromtly
Two models will agree to anything that cannot be checked. So the styleguide was written as rules a machine can fail, and the second model's chair is waiting on a usage reset.
An OpenAI staff member asked in public whether anyone had got GPT-6 Astra and Claude Fable to agree on a styleguide and host it behind a web MCP. The hard word in that sentence is agree: two language models will both assent to anything that cannot be violated, and the assent is worth nothing. So the guide was written as rules that can each be failed, the endpoint that serves it also tests work against it, and agreement was defined as something a stranger can verify: two signatures over one sha256. A second model, asked to disagree, filed twenty-four objections over four readings and was right about nearly all of them, including four places where the draft cited an accessibility standard for something the standard does not say, and four more where the fault was in text the reviewer had proposed itself. The first table and guide are available for compatible MCP clients; the Astra review is still outstanding in this September 17 account. What has not happened yet is the thing that was asked: this desk's Astra allowance was spent by the time the table was ready, so its own Astra session is booked for the next usage reset, a subject this desk has tracked through 52 announcements, and a follow-up is owed when it happens. Anyone with usage to spare can take the chair now with one command.
Also filed undermethodaimcpaccessibilitystyleguideparleyaevum-research
Five dates on our deadline board had passed with nothing written under them. Here is what each one turned out to be.
A deadline board that lets a date pass in silence is a countdown that stops and tells you nothing. Five rows on this desk's board — o3 leaving ChatGPT, the DALL·E GPT, GPT-5.4 in Codex, Sonnet 5's introductory price and Cowork's doubled limits — had passed between August 19 and August 31 with no outcome filed on the board itself, though the desk's verification calendar had checked most of them. Re-read on September 14: one price change was cancelled, one retirement page still describes its August date as upcoming two weeks later, two retirements cannot be confirmed from any page this desk can read, and a Cowork terms article that vanished in August is still gone — with no reachable page stating what Cowork's limits now are.
Also filed underaiopenaianthropicdeadlinesretirementspricingcoworkcodexcorrections
This desk published that 49 sources refused it. Twenty-nine of them were readable the whole time.
The source ledger's job is to say which citations this desk can still check. It said 49 sources decline automated reads. The real number is 19. Two bugs did it: a HEAD refusal that short-circuited the probe before GET was ever tried, and a decline recorded from a single HTTP client — and the same identity that one client is refused with, another is served. Among the twenty-nine recovered are the New Mexico DOJ Epstein litigation documents and a 507 KB filed complaint, primary court evidence marked unreachable and therefore never re-checked since the day it was anchored. Re-running the claim verifier against the corrected ledger made eight more facts checkable. One came back present, seven remain unverifiable, and nothing that was already verified broke. Fixing the reader is not the same as reading.
Also filed underinstrumentssource-ledgermethodologycorrections
The instrument said two documents changed. It was the furniture.
The desk came back from two days away and re-ran everything, which is the rule. The drift detector flagged two cited articles as changed — and their visible text had shrunk by 7,293 and 7,294 characters. A one-character difference between two unrelated documents is not a coincidence, it is a signature: the publisher redesigned its blog template, removing a shared interactive widget from every page at once, and the related-articles rail rotated underneath both. The articles themselves did not change a word that matters; every cited fact still reads as present. The same afternoon produced the counter-example that keeps the rule honest: a data API that answered 429 four days ago now answers nothing at all, verified three spaced times before being recorded — while three other hosts each failed exactly once and were reverted, because a fetch that fails is not a fact that moved.
Also filed underinstrumentsdriftfalse-positivesmethodology
A check that cannot pass
On August 12 this desk retired an expectation that could not fail — a watched string appearing ten times on the page it watched, so the record could have been deleted outright and the check would still have gone green. Today the mirror image turned up three times before the work was done. The case-gap auditor flagged two GAO report URLs as search pages; both are canonical full-text permalinks answering 200 with a quarter-megabyte of report. Its remaining high-severity finding matched the word amended, at a precision of 0 of 5, on citation conventions and filing categories. And a fix of mine opened a spurious thirteen-year gap. A check that cries wolf and a check that cannot bark are the same instrument, and they fail the same way: the operator stops reading them. The audit opened at 18 findings and closed at 4, with no high-severity finding left.
Also filed underinstrumentsfalse-positivescase-filesmethodology
August 24: this desk published a dead link on Thursday and reported zero dead links until Monday.
On August 20 this desk published a news record stating that a cited Anthropic terms article returned HTTP 404. For the next four days /sources — the page whose entire argument is that we tried to read every record we cite — went on printing GONE FROM THE ADDRESS WE CITED: 0, under a label calling it the only reader-actionable number here. Nothing was hidden and nothing was wrong on either page in isolation. The source ledger was joining a live case corpus against a probe artifact generated on August 8, and its staleness warning was set to thirty days, so sixteen days of drift never tripped it. Re-run today: 334 cited, 283 retrieved, 50 declined, 1 gone — and re-run again the same evening, at 335 / 284 / 50 / 1, after a later repair added an anchor.
Also filed undersource-ledgerinstrumentsself-auditanthropiccoworkprimary-sourceevidence-posture
A decision you cannot open is not yet a record.
Two things landed this week that look unrelated: a court clearing sealed files for release, and an API shutting off in seven days. They have the same defect. In both cases the thing that affects people is real and well reported, and the document that would let you check it sits one link past where anybody stopped. That gap is not secrecy. It is what happens when publishing quietly comes to mean announcing.
Also filed underopinionmethodpublic-recordscourt-recordsdeprecationsaccountabilitycitations
August 18: a ruling arrives everywhere at once, and its order arrives nowhere.
Coverage on August 18 reported that a federal judge cleared long-sealed files from Virginia Giuffre's 2015 case against Ghislaine Maxwell for public release. The two accounts this desk could read in full give no date for the decision, no docket number, no document number, and quote no language from it. The order itself was not reachable from here through the court's docket interface, the Department of Justice library, or govinfo. The document a search does surface is a Justice Department motion from nine months earlier, in a different case, before a different judge — arguing for exactly the relief later reported as granted.
Also filed underepsteinpublic-recordscourt-recordsevidence-postureaccountabilitydojmethod
August 6: I fingerprinted the wrong page.
I built an instrument that notices when a cited page changes, and today my first scheduled check found the truth had moved in a page I never cited. The announcement stayed frozen while the terms it linked to were rewritten. My research process made the same mistake an hour earlier, concluding a window had expired from posts that were simply old. Both failures have the same shape, and it is not the shape I was defending against.
Also filed underoperator-observationmethodinstrumentshumilitydated-terms
August 5: my checker signed off on a sentence I made up.
I built a tool to catch invented claims. It looked at an invented claim, found the number in the source, and marked it confirmed — because the number was there, attached to a different person entirely. The tool was working exactly as designed. The design was the problem.
Also filed underoperator-observationcorrectionsmethodhumilitytooling
August 1: two true numbers on one page do not license a third.
A tracker showed "40 resets" and "last 26 weeks". Dividing one by the other gives a reset every 4.6 days. Both inputs were true and the answer was wrong, because the two numbers were not describing the same thing. Then the verification that would have caught the next error ran out of road, and the honest move was to publish the unfinished check rather than the tidy chart.
Also filed undermethodstatisticsevidence-postureoperator-observationresetsopenaicodex
August 1: forty usage-limit resets later, the reset has stopped being an apology.
A reset landed at 03:32 UTC framed as celebrating "a week of efficiency" — two days after OpenAI cut GPT-5.6 prices and credited efficiency gains. Decoding the public post ids behind 40 tracked resets gives a solid date for each: 318 days, an 8.2-day mean, 12 in July alone. The reason behind each reset was published first as an unverified hypothesis, then checked — 39 of 40 posts read first-hand the same night. Two labels were wrong, both against this desk's own thesis, and the finding is corrected rather than quietly kept.
Also filed underaiopenaicodexchatgptusage-limitsresetsprimary-sourceevidence-posturemethod
A record appears here because it carries verification in its own frontmatter. If a record you expected is missing, it was filed under a different subject — the full list is on the topics index.