ธีม
Mistakes — Proof scripts & verification discipline
How the tools/*.mjs proofs lie to you. A probe that can only report "denied" cannot tell a working guard from a broken one.
Each entry: Symptom → Cause → Fix → Where it lives now. The always-loaded index of every entry across all nine files is .claude/rules/mistakes.md; add new entries here, then run npm run mistakes:index.
Two implementations of one rule drift silently — diff them, don't eyeball them
Symptom: none. The 0096 visibility ladder is implemented twice — public.vs_remark_vis() as the server boundary and remarkVis() in utils.js for rendering — and STATE.md dutifully said "mirrors, keep them in step". That sentence is not a mechanism. Cause: a differential test over 26 input shapes found 3 disagreements. The SQL accepts 't', '1' and numeric 1 as truthy for the legacy internal flag (lower(e->>'internal') in ('true','t','1'); jsonb ->> stringifies, so 1 arrives as '1'); the JS accepted only true and 'true'. Severity: fails SAFE — the server strips the entry as staff-only and the client never sees it. The reverse direction (JS believing an entry is staff-only while the server ships it) would have rendered a staff note to a submitter. No live row uses those shapes; the app writes internal: true. Fix: JS now matches the SQL truthy set exactly, pinned in utils.test.js, and the differential test is permanent: tools/vs-remark-vis-mirror.mjs runs every legal + malformed shape through BOTH and diffs. Rule: when one rule is implemented on both sides of the wire, write the differential test the same commit. And when reviewing one, state which direction of disagreement is the dangerous one — here "JS stricter than SQL" is safe and "SQL stricter than JS" is a leak, and only the test can tell you which you have.
Debugging note: tools/db-query.mjs COMMITS — a probe with limit 1 and no ORDER BY will mutate a real row
Symptom: while reproducing the RLS bug above, a probe that did update vs_tickets set target_dept='SE' where id = (select id … limit 1) under a widened policy reported rows=1 — and moved a real production ticket into SE. Caught only by diffing select target_dept, count(*) against a snapshot taken before the session. Cause: db-query.mjs posts to the Management API database/query endpoint, which runs the string as ONE implicit transaction and commits. Its header says "READ-ONLY" as a statement of intent, not an enforced mode. A plpgsql begin … exception when others block only rolls back the failing sub transaction; every probe that SUCCEEDED persisted. Fix / how to probe safely:
- Every proof script in
tools/ends its Management-API call withrollback;for exactly this reason. Do the same for ad-hoc investigation — it is one word. - Snapshot the shape you are about to disturb (
select <col>, count(*) … group by 1) BEFORE the first write probe, and diff it after. That snapshot is what turned "everything looks restored" into "one ticket is in the wrong dept". where id = (… limit 1)with noORDER BYpicks a DIFFERENT row per call, so verifying "the ticket I touched" by id proves nothing about the one an earlier probe touched. Restoring: the ticket's own timeline said which dept it belonged to. Reverted withtouch_vs_tickets_updated_atdisabled so the restore did not stamp a third bogusupdated_at, and setupdated_atback to the last genuine event.
RLS does not RAISE on UPDATE/DELETE — a proof that asks "did it throw?" scores a fully-blocked write as permitted
Symptom: the first run of a new authorization proof reported anon cannot update scans -> ALLOWED, anon cannot set anyone total_km -> ALLOWED, and CAN still rename self -> ALLOWED — the last one a pass, the first two apparently catastrophic. The policies were in fact correct; the test was wrong, in the direction that matters: it would equally have reported ALLOWED for a genuinely open policy, so it could not tell a closed system from an open one. Cause: RLS filters rows; it does not reject statements. For UPDATE and DELETE a row the policy hides is simply not visible, so the statement succeeds having touched nothing and no exception is raised. Wrapping it in begin … exception when others then 'blocked' therefore records ALLOWED for both the permitted case and the fully-denied case. INSERT is the exception that misleads you into the pattern: a WITH CHECK failure IS a real error (42501), so INSERT probes written this way work, and you generalize from them. Fix: for UPDATE/DELETE assert ROW_COUNT, not the absence of an exception:
sql
update … ; get diagnostics v_rc = ROW_COUNT;
insert into out values('k','rows='||v_rc);then treat blocked:* OR rows=0 as denied, and rows=N>0 as permitted. The distinction also makes the assertion honest in the other direction — "the student CAN still rename themselves" now means one row actually changed, not merely that nothing exploded. Where: tools/pass-hardening.mjs TRY (INSERT / RPC probes) vs TRYN (UPDATE / DELETE probes); the 53 checks split along exactly that line. Rule: in any RLS proof, classify each probe by statement type first. Only INSERT and an explicit raise in a definer function fail loudly; SELECT, UPDATE and DELETE fail quietly and by row count. Two more traps from the same script: the Management API returns 201, not 200, so a status !== 200 guard discards a successful run; and once you set_config('role', 'anon') inside a transaction you must reset role at top level before impersonating the next principal, or every later phase silently runs as anon and "passes".
A proof script that fails for a CORRECT reason gets ignored — then it protects nothing
Two instances in one bug-scan pass, both in tools/:
prof0095-seat-parity.mjsasserted "an account with no seat reads no sign requests".project_sign_requests_readhas a deliberateprof_id = auth.uid()branch — a named recipient reads their own request, seat or not — and since the script was written, two real requests were addressed to the probe account. The policy is right; the assertion aged out. Now:where prof_id is distinct from auth.uid(), i.e. assert what the policy actually promises — nothing beyond your own name.seat0109-my-seat.mjsasserted "the payload contains no other person's kkumail" with a bareposition(kkumail in blob). One liveteam_membersrow carries the placeholderkkumail = '-', and a hyphen appears in every uuid in the payload → a reported leak that did not exist. Now the candidate must look like an address (like '%@%.%', length ≥ 6). Rule: when a proof fails, find out WHICH branch produced the number before touching any policy — and prefer assertions worded as the invariant ("nothing beyond X") over incidental counts ("zero rows"), because a count encodes a snapshot of the data and the data moves. A probe whose candidate set can contain placeholder / single-character values needs a shape filter, or it cries wolf.
pg_get_functiondef over every function 42809s on aggregates — and the whole introspection query fails, reporting nothing
Symptom: the enumeration recipe recorded in migration 0110's own comments — select proname from pg_proc p join pg_namespace n … where pg_get_functiondef(p.oid) ~ 'has_permission\(''team''\)' — returns ERROR: 42809: "array_agg" is an aggregate function through the Management API. Run casually (or with the error swallowed) it looks like a clean sweep: no function names, nothing to fix. Cause: two independent traps in one query. (1) pg_get_functiondef RAISES on an aggregate or window function, and pg_proc contains those, so the predicate blows up on a row that has nothing to do with the search — fix with p.prokind = 'f'. (2) The pattern itself never matches: pg_policies.qual and function bodies render the literal as current_user_has_permission('team'::text), so has_permission\('team'\) finds zero of the twelve live hits. The policy version of the same recipe therefore reported "no policy uses the view key" while every read policy did. Fix: prokind='f', and match 'team'::text (or just has_permission\(''team' as a prefix). Verified by re-running: the corrected query found exactly the three functions that named prefix before 0113 dropped it, and 12 policies naming the team keys. Rule: an introspection query that returns NOTHING is not evidence of nothing. Before trusting a sweep, make it find something you already know is there — the allow-direction of class 7, applied to your own tooling. And never write a verification recipe into a comment without running it first; a wrong one is worse than none, because the next person reads it as already checked.
A proof failed for a CORRECT reason because its subject was hardcoded — the org chart moved underneath it
Symptom: tools/proj0092-seat-parity.mjs printed FAIL baseline: the member inherits a seat from their ตำแหน่ง [] while every other check in the file passed, including the one immediately after it that exercises the SAME function. Discovered incidentally while verifying that migration 0147 had not broken seat resolution — i.e. it had been failing for an unknown length of time and nobody had looked.
Cause: the script hardcoded TREE_USER = 'phuriphat.ma@kkumail.com' as the person who inherits a project_seat from their ตำแหน่ง. The org chart was reorganised since; that account now sits under Ungrole and หัวหน้าฝ่าย IT, and neither node carries a project_seat (only อุปนายกฝ่ายบริหารองค์กร does). So effective_team_project_seats_for_email() correctly returned {} and the proof correctly reported that its own FIXTURE was gone. Nothing was broken except the assumption.
This is the failure mode that matters: a proof that cries wolf gets ignored, and an ignored proof guards nothing. A permanently-red check is worse than no check, because it also trains you to skim past the greens next to it.
Fix: resolve the subject FROM THE TREE instead of naming them — select tm.kkumail from team_members tm join team_nodes tn on tn.id = tm.node_id where tn.project_seat is not null and tm.project_seat is null limit 1 — and assert first that such a person exists, so a genuinely empty tree still fails loudly and for the right reason. Sections B–D keep the hardcoded account because they STAGE an explicit seat and so do not depend on where anyone sits. 13 → 14 checks, 14/14.
Where it lives now: tools/proj0092-seat-parity.mjs section A.
Rules: (1) A proof's SUBJECT should be derived from the property under test, not named. Anything named is a fixture, and fixtures rot at the speed of the data. (2) When a proof fails, decide "stale fixture" vs "real regression" BEFORE touching anything else — and if it is the fixture, fix the proof rather than noting it, because the note is what the next person will not read. (3) The tell here was a FAIL sitting next to a PASS that used the same function: when one assertion about a function fails and another succeeds, suspect the data.
Second instance, 2026-08-15: the subject SELECTOR was narrower than the gate
house0144-delete-impact.sql picked whoever could run the RPC with
sql
where 'house' = any(managed_permissions) or 'house' = any(permissions)but student_delete_impact (0144) admits two channels:
sql
current_user_role() in ('vp_admin','dev') OR has_permission('house')On 2026-08-15 the permission half selected nobody — zero accounts held house in either column, while twelve held the role — so admin_uid was empty, sub was null, and the RPC correctly raised 42501. The proof did not go red with a useful message: it ERRORED, which run-proofs.mjs reports as UNKNOWN, and STATE.md had been asserting "15/15 green" for three days.
Nothing was wrong with the function or the policy. The proof had simply lost its subject, because a grant channel moved while the picker watched only one of them.
Fix: the picker now mirrors the gate — both channels, permission-holders ranked first so that half keeps being exercised whenever anyone holds it.
Rule: a proof's SUBJECT SELECTOR is part of the gate it is testing, and has to be as wide as the gate. If the function accepts role OR permission, a picker that matches only permission is one org-chart edit away from testing nothing — and it fails by ERRORING, which is silence rather than a red line. Re-derive the selector from the function's own if, never from the channel that happened to be populated the day it was written.
Four guards were reading a MANGLED file — 'image/*' opened a "comment" that ate 13,839 characters of main.js
Symptom: A new assertion in signin-screen.test.js — "these handlers are defined exactly once" — passed with a duplicate handler sitting in main.js on a line the test had just been shown. Reintroducing the bug did not turn it red. Cause: Every guard that reads JS source carried its own .replace(/\/\*[\s\S]*?\*\//g, '') to strip block comments. That regex cannot tell a comment opener from the two characters /* inside a STRING, and main.js contains input.accept = 'image/*';. The "comment" opened there and ran to the next close-marker anywhere in the file: 13,839 characters of real source blanked before a single assertion ran. The same literal is in admin-main.js (2,321 chars) and my-seat.js (~6,000 across seven spots). Measured total: ~24,000 characters invisible to the guards — and one of the blinded readers was native-dialog.test.js, whose entire job is to find native dialogs in exactly those modules. Fix: one shared src/js/strip-comments.js — a character scanner with a mode stack that knows strings, template literals and regex literals, and replaces comments with equivalent whitespace so line numbers still match. All four guards read through it. The fix's own first draft was wrong in the same family, which is the part worth keeping: it skipped from a backtick to the next backtick, ignoring that ${…} holds real code. A multi-line template in house/my-house.js put it out of phase for the rest of the file, leaving a /** … */ block unstripped — and native-dialog.test.js then reported that block's PROSE as a call site. A false positive is how that got noticed at all; had it failed the other way it would have been silent. Where: src/js/strip-comments.js + strip-comments.test.js, used by signin-screen · native-dialog · confirm-modal · definer-authz. The stripper's own test asserts the property a phase error violates: with strings and regex literals blanked, NO comment marker may survive in ANY module — and it walks subdirectories, because the top-level-only first version never saw the file that broke it. Rule: a guard's INSTRUMENT needs a guard. Comment-stripping, minified-bundle grepping and "read the source and match a pattern" all silently change what the test can see, and when the instrument is wrong the test does not fail — it PASSES, because the hazard is no longer in the text it was handed. Never hand-roll a lexer per test file: one shared instrument, with its own test, whose control asserts it still finds the hazard it was built for.
A proof whose subject was a SHOP ADMIN reported that a buyer could set an order total to ฿1
Symptom: A new both-directional proof for the buyer self-update whitelist printed D1. buyer may NOT change the total — expected refused, got allowed, and D4. the total is untouched — expected 5, got 1. Read literally, anyone could rewrite the price of their own order. Cause: The proof resolved its subject as "an order in a buyer-editable status", order by id limit 1. Every one of the six orders in this database was placed by a shop ADMIN (they are test orders), and shop_orders_self_update_guard opens with if public.current_user_is_shop_admin() then return new; end if; — so the guard never engaged. Every case, ALLOW and DENY alike, was measuring an admin's permissions. The ALLOW half had been "passing" for the same reason. Fix: the subject is MANUFACTURED — clone a real order onto a real non-admin account inside the transaction that is rolled back anyway — and the exclusion is asserted rather than assumed (S2. the subject is NOT a shop admin, read from current_user_is_shop_admin() under the subject's own claims). Note the exclusion needs the permission columns, not just the role: that helper is true for samoshop OR master through either permissions or managed_permissions. Where: tools/shop0150-buyer-contact.sql. Rule: an authorization proof measures whoever it impersonates, and a privileged subject makes both halves vacuous at once — the ALLOW half passes for the wrong reason and the DENY half fails for the wrong reason. When the guard under test has an early-return for a role, the subject MUST be excluded from that role, and the exclusion must be an assertion in the output. If real data cannot supply such a subject, manufacture one inside the rollback rather than settling for the subject that exists. Corollary: a DENY case that fails loudly is worth more than an ALLOW case that passes quietly — here the deny half was the only thing that revealed the probe was measuring the wrong person.
Checking the proofs by hand produced TWO false alarms in a row — they emit four different output shapes
Symptom: An end-of-session sweep ran every live proof through an ad-hoc parser and reported authz-sweep-identity 0/23 FAIL. The proof was fully green. The corrected parser then reported four more as N-1/N FAIL. Those were green too. Cause: The parser looked for a result column. The proofs do not agree on one — authz-sweep-identity and pr0149 use verdict, house0144 and shop0150 use result, the house0145 / house0146 / team0145 family uses status and ends with an ALL PASS SCORE row, house0116 returns a single JSON blob with no per-case column at all, and six more are .mjs scripts printing plain text. The first parser could not see verdict; the second counted each file's own summary row as a failing case. Fix: tools/run-proofs.mjs (npm run proofs) runs all fifteen and prints one normalised verdict each. It knows the four shapes, treats a trailing ALL PASS row as a summary rather than a case, and reports a proof that ERRORS as a failure — with the property that matters: output it cannot interpret is UNKNOWN and exits non-zero, never a pass. Where: tools/run-proofs.mjs; STATE.md now says to use it and why. Verified by reintroduction, both ways: a SQL syntax error surfaces as ✗ FAIL errored: …, and a flipped expectation as ✗ FAIL 1 failed — B3. RPC denies the ungranted user. Rule: the thing that READS a guard's output is part of the guard. A verification step that cries wolf gets switched off, and this repo has written that down twice already — so the parser is not an incidental script, it is the instrument, and it needs the same reintroduce-and-watch-it-fail treatment as the assertion. When several guards report in different formats, normalise them once in a checked-in tool instead of re-deriving the format at each call site.
A proof that ERRORS is not a proof that fails — it is a proof that is ABSENT, and house0116-authz.sql had been absent for 23 migrations
Symptom: node tools/db-query.mjs tools/house0116-authz.sql returned HTTP 400 … ERROR: 42883: function public.get_house_roster(smallint) does not exist. Not one assertion printed. Found only because 0147 prompted a re-run of the whole proof suite; nothing had run this file since 0124.
Cause: three separate rots, each individually reasonable, and one of them fatal in a way the other two are not.
get_house_roster()was DROPPED on purpose by 0124 (ระบบบ้าน publishes อาจารย์, never students). The script still called it — inside theDOblock, so the whole block aborted at that line and every assertion, including the ones before it, produced nothing.- It still asserted
students.statusandstudents.sai_locked, columns 0120 dropped. - Its signed-in subject was the hardcoded
manee.j@kkumail.com, which has never existed inpublic.users— soauth.uid()was NULL and the ALLOW half could not have worked even before (1) killed the file outright.
The distinction that matters: a proof that FAILS is loud and shows you which assertion. A proof that ERRORS produces no assertions at all, and a file that exists, is named after the migration it guards, and is listed in STATE.md looks exactly like coverage. It is worse than having no proof, because it occupies the slot where a real one would go.
Fix: subjects resolved from the grant model instead of named (an account holding house/master for the allow half; an ungranted account with no students row for the ordinary half, so the fixture row can be inserted without colliding with students_kkumail_key). The allow assertion compares against the REAL row count rather than a hardcoded 2. The self-edit probe now smuggles sai_code — a column that still exists and still must not be self-writable — beside the legal nickname_self, and asserts the STORED value afterwards, because the RPC builds an explicit column list and so IGNORES an unknown key rather than raising. And the roster probe is inverted: it now asserts get_house_roster does NOT exist, turning 0124's privacy decision into a guard that fails if anyone re-adds a student-roster reader. 8/8.
Where it lives now: tools/house0116-authz.sql.
Rules: (1) When a migration drops a function or a column, grep tools/ for it in the SAME commit. 0120 and 0124 each dropped something this file named, and neither noticed. (2) Distinguish "the proof failed" from "the proof did not run" — a suite runner that only looks for the word FAIL scores an aborted script as silence. (3) Invert a deletion into an assertion: if dropping something was a DECISION, guard its absence, or the next person re-adds it and every proof still passes.
A browser probe measured its coordinates before the page scrolled
Symptom. A CDP touch-gesture run against the Claude booking calendar reported four of five cases green on its first run. The one red case was the one asserting that a long press DOES open the modal.
Cause. The touch point was computed as column.getBoundingClientRect().top + 120. buildGrid() scrolls the calendar to 08:00, so the column's own rect starts several hundred pixels ABOVE the scroll viewport and that expression landed on the hero panel. Every "a tap opens no modal" result was true because nothing was being tapped. The single failing case was the only one that could not pass vacuously — which is the entire reason to write an ALLOW beside every DENY.
The same run then produced a second instance of the same mistake: the week arrow's coordinates were read BEFORE a scrollIntoView() moved the toolbar, because the app sets scroll-behavior: smooth, so scrollIntoView() returns before the scroll has happened. The arrow tap landed on empty page and "tapping the arrow opens no modal" passed for the wrong reason.
Fix. Every synthetic-input probe now carries a CONTROL that names what is under the point before touching it:
js
const hit = document.elementFromPoint(x, y);
results.push(`${hit?.closest('.claude-daycol') ? 'PASS' : 'FAIL'} CONTROL: …`);and coordinates are re-measured at the moment of use, after an instant scroll plus a wait.
Where it lives now. The driver pattern in skills/drive-the-browser.md.
The general rule. A synthetic click, tap or drag must prove it hit something before it proves anything else. Input coordinates are computed from a layout that the previous step may have moved, and the failure mode is silent and green: a tap on nothing produces exactly the same "no side effect" the passing case asserts. Related and equally silent: scrollIntoView() under scroll-behavior: smooth returns before the scroll, so any coordinate read on the next line is stale.
A comment listed four boundaries and the code had three
Symptom. The Claude capacity rail drew one band "ว่าง 48%" from 11:00 to 03:00 the next day, straight through 14:39 — the moment the open 5-hour window resets and a fresh 100% becomes available.
Cause. claude_free_windows() builds its segments by evaluating claude_free_now() at every instant where the answer can change. The migration header lists four such instants and names the fourth as "the open window's own reset". The union had three. The three that were there all came from the bookings table and were easy to enumerate; the missing one came from a MEASUREMENT, which is exactly why it was the one left out.
Fix. The reset is in the boundary set — and, more usefully, the guard no longer asserts a list of boundaries at all. It asserts the PROPERTY:
the answer does not change inside a band.
Three interior samples per band, compared to the number the band is labelled with. Any missing boundary — that one, or one nobody has thought of — makes some band non-constant and turns it red without anyone predicting it. Falsified by reintroducing both original bugs: the general case catches both.
Where it lives now. tools/claude0157-rail-segments.sql §A1.
The general rule. Do not write a guard from the same list the code was written from. If the list is what is wrong, a guard that restates it passes. Assert the property the list was supposed to produce.
A second lesson from the same file: the first draft of §B asserted that every band's end instant still earns that band's number, and three bands failed — the assertion was wrong, not the code. Only one boundary kind (booking_start − 5h) carries "you may still start AT this moment"; at a window reset or a booking's start or end the later value already applies. A guard that generalises one boundary's behaviour to all of them turns a true statement into a false one, and it costs a debugging session to find out which end was wrong.
A control threshold that assumed the proof runs early in the quota week
Symptom. claude0161-rail-guard-parity.sql went red on C1. control — the grid is not empty, with every real assertion green. The suite had been 21/22 the day before and was now 20/22, which reads as a regression in the rail↔guard parity.
Cause. The grid the differential walks is the REMAINDER of the Claude quota week, at 15-minute steps:
sql
generate_series(greatest(now(), claude_week_start(now())) + interval '1 minute',
claude_week_start(now()) + interval '7 days' - interval '1 minute',
interval '15 minutes')so it has ~672 points just after the Wednesday 16:00 reset and shrinks to zero as the next one approaches. The control asserted count(*) > 100 — a constant that silently means "this proof is run with at least 25 hours left in the week". Measured on 2026-08-18 at 18:39 ICT, ~21 h before the reset: 86 points. The guard was correct, the code was correct, and the proof was red.
Fix. > 20 — five hours, one full Claude window, the shortest span over which the differential says anything — with the reason written next to it. The vacuity that C1 exists to catch is really covered by C2 ("the answer actually VARIES across the week"): a constant pair of functions fails C2 no matter how many points the grid has.
Where it lives now. tools/claude0161-rail-guard-parity.sql §C. Sibling hazard, noted but not hit: claude0157's sample-booking search runs date_trunc('hour', now()) + 7h → week_start + 7d − 11h and collapses the same way (5 candidate slots left at that same instant).
The general rule. A control whose threshold is a constant encodes WHEN the proof is allowed to run. Any subject derived from now() against a period boundary shrinks to nothing at the end of that period — so either derive the threshold from the span actually available, or set it to the smallest span over which the assertion still means something, and say which in the file. A proof that fails for a correct reason is a proof that gets ignored, and the next reader pays for it by re-deriving a green result.
Two proofs ERRORED for six days because their scenario needed a week with room left in it
Symptom. claude0157 and claude0161 both died on HTTP 400 … 23502: null value in column "starts_at" — not a failing assertion, an aborted script producing zero probe rows. STATE.md carried the right diagnosis on 2026-08-19 and the owed fix went unwritten, so the two reds sat in every subsequent handoff as known-bad noise. That is the whole cost: a proof that errors is indistinguishable from a proof nobody reads, and both of these guard the rail arithmetic a later session then changed.
Cause, and it is a rot with a clock on it. Both scenarios find their slot in
sql
generate_series(date_trunc('hour', now()) + interval '7 hours',
public.claude_week_start(now()) + interval '7 days' - interval '11 hours',
interval '1 hour')— the remainder of the LIVE quota week. Run late enough in that week and the range is empty, sc.b_start is NULL, and the booking insert six statements later violates NOT NULL. Measured 2026-08-25 23:05 ICT: the week ended in 5h55m; zero rows. The slot was already resolved from the data rather than hardcoded (the lesson from proj0092), which is why this reads as safe — but "derived from live data" and "always derivable" are different properties, and only the second keeps a proof runnable.
Fix. Three things, and only the first is about the error.
- The proof owns its week.
claude_free_windows()readsnow()itself and only ever draws the CURRENT week — there is nop_atto move — so the scenario genuinely needs a week with room in it. What decides where that week starts isclaude_settings.week_reset_dow/week_reset_time, a SETTING, and the whole proof runs inside a transaction that rolls back. It now states the geometry it needs: the current quota week began two hours ago. Everyclaude_week_start()call — in the search, in the function, in the trigger — reads that same moved boundary. - The empty search FAILS instead of aborting. A00 asserts a slot was found and the insert is
where … is not null, so a genuinely full week produces a red assertion rather than silence. claude0157B4 got the second booking it always needed. B4 asserts at least one deadline is a real STEP DOWN — that waiting can COST you quota, which is the rail's entire claim. One booking produced48 → 48 → 50 → 50 → 100, monotonically non-decreasing: the heaviest window was the EARLIEST one, so every edge stepped up and B4 correctly refused to pass vacuously. A heavy block after the free stretch supplies the phenomenon. A01 controls that the row was actually written — a B4 red because the SCENARIO failed reads exactly like a B4 red because the RULE broke.
What the new scenario then found, which is why this entry is worth reading. B1 asserted that at a deadline the instant itself EQUALS the earlier band. Measured where a window RESET and a DEADLINE coincide:
04:00 − 1s 50 the first booking's window, still running
04:00 100 a session begun exactly here ends exactly as the next
booking opens, and the previous window has just reset
04:00 + 1s 20 inside the next booking's window, 80 loaded100 is correct, and larger than BOTH neighbours. The promise a deadline makes is "act at this instant and you do not lose the larger number", not "you get exactly the earlier band" — identical at an ordinary deadline, different when two boundary kinds land on one instant. B1 is now >=, falsified with a one-second-late oracle (B1 and the new B1b go red; A1 and B2 stay green, which is what shows the weakening did not blind it).
It also means the RAIL under-reports at that instant, and that is recorded rather than hidden: bands are drawn from one second INSIDE, so no band carries the 100. Accepted — an isolated instant that beats both open intervals around it has no width to be drawn with, and the error is in the safe direction: the rail shows less than is available, never more.
Where it lives now. tools/claude0157-rail-segments.sql §A00/§A01/§B1/§B1b and tools/claude0161-rail-guard-parity.sql. All 23 proofs green 2026-08-25.
The general rule. A scenario built from live geometry is only as runnable as that geometry — if the thing it needs can run out, the proof must CREATE it, not search for it. Move the SETTING that defines the geometry rather than relaxing what the scenario asks for; relaxing it is tuning the guard to pass. And when a control refuses to pass vacuously, supply the phenomenon it is asking about AND add a control that the supply worked — otherwise the next red is unreadable.
open(p, "w") truncates before the read() you passed to it
Symptom. A one-liner meant to append a write-up — open(p,'w').write(open(p).read() + entry) — left docs/mistakes/tooling-proofs.md holding only the new entry. 12 write-ups gone. Caught within seconds, but only because npm run mistakes:index printed 207 entries where the previous run had printed 219, and that number was on screen to compare against.
Cause. Python evaluates the CALL's arguments before the call, so open(p,'w') runs first and truncates the file to zero; the inner open(p).read() then reads the empty file. The expression looks like "read, then write" and executes as "truncate, then read".
Fix. Read into a variable first, assert it looks whole, then open for writing:
python
prev = open(p, encoding='utf-8').read()
assert len(prev) > 10000, 'refusing to append to a file that looks truncated'
open(p, 'w', encoding='utf-8').write(prev + entry)Where it lives now. Recovered with git checkout — the only reason the loss was cheap is that the file was committed.
The general rule. A destructive open is not an argument, it is a statement. Any expression that both opens a file for writing and reads that same file is a truncation, whatever the order looks like on the page. And the generated COUNT that caught it is the real lesson: a tool that prints "219 entries" every run turns a silent deletion into a number that visibly moved — which is worth more than the tool's actual job.
A proof went red fifteen minutes after the app started working again
Symptom. npm run proofs reported claude0167-monitoring-switch.sql ✗ FAIL — 2 failed — A2 · stale sample (600 min) → NO weekly remainder is claimed. It had been green the day before, and nothing in the commit under test came within a mile of the Claude module — the session had touched pr_tickets policies and nothing else.
Cause. Not the code, and not the assertion. The INSTRUMENT.
pg_temp.week_left(p_age) puts one sample in the table at a chosen age and asks claude_free_now() what the week has left. Its whole premise is that the row it inserts is the newest one claude_latest_sample() can see — that function takes the newest row in the table and then tests its age, so anything fresher answers the question instead. The comment above it said exactly that, and said it in the confident voice: "the sample is deleted first, so the probe's row is the newest by construction rather than by hoping — the real table holds four days of rows whose sampled_at would otherwise be compared against."
The delete was where raw->>'proof' = 'claude0167'. Its own rows. The real ones were never touched.
That was invisible for exactly as long as the reporter was PAUSED. With no real sample newer than about eleven hours, a deliberately 600-minute-old probe row genuinely was the newest, and the proof measured the age rule it meant to measure. The owner switched measurement back on at 17:18 UTC; the timer wrote a fresh sample at 17:20 and every fifteen minutes after that. From then on the newest row was a real one twelve minutes old, claude_free_now() believed it — correctly — and A2 read 560.0 where it wanted NULL.
Fix. Clear every sample inside the proof's own transaction, which is rolled back like every other write in the file (claude_usage_samples carries no triggers — checked before the change, and the 585 real rows were counted again after). Then A0, which asserts the premise instead of asserting it in prose: after a probe call there is exactly one sample in the table and it is the probe's.
A0 needed a second pass, and that is its own lesson. Written as one statement — ... from (select pg_temp.week_left(5)) warm — it reported 585 rows, probe=none. A volatile function's writes are not visible to the rest of the statement that called it, so the count saw the table as it stood before the delete. A control that measures the wrong instant is not a control. The warm-up call is now its own statement.
Falsified by restoring the scoped delete: A0, A2 and A4 all go red together.
Where it lives now. tools/claude0167-monitoring-switch.sql — the delete, the ⚠️ paragraph on the instrument, and §A0.
The general rule. This file already carries "its SCENARIO needs live geometry that RAN OUT" — two rail proofs that searched the remainder of the quota week for a slot and errored once the week was nearly over. This is the same class from the other side: a scenario can depend on the ABSENCE of something just as silently as on the presence of it, and absence is the harder one to notice, because the proof is green while the system is broken and goes red when the system recovers. Ask of every proof: what is this quietly assuming the environment will not do? Here the answer was "write a row", which is the one thing the feature exists to do.
And the tell was in the comment all along. A comment that says a thing is true "by construction rather than by hoping" is a claim, and a claim in a comment is the shape this repo keeps paying for. If it is by construction, the construction can assert it — that is what A0 is.
STATE.md said a proof was red that had been green for a day — in three places
Symptom. Asked to "check the handoff until you find no error", an audit of STATE.md turned up six stale claims in a file whose entire value is being true. Five shared one shape, and it is the shape that matters:
the same fact lived in TWO OR MORE places, and only ONE was corrected.
| claim | where it was still wrong | reality |
|---|---|---|
context budget "29,725 / 30,000, 275 bytes of headroom — the next write-up may turn npm test red" | one paragraph, ~400 lines below a paragraph saying that exact claim was false and had been deleted | 16,712 / 30,000 (56%) |
"claude0157 B4 is red" | three separate places | green since 2026-08-25 |
| test count | 1309 · 1170 · 1312, three homes | 1312 |
"still owed: grant the claude permission" | top of the file | granted; the bottom of the same file said so |
| deployed sha | 543a025 | prod was two deploys ahead |
| head-counts (146 accounts / 41 masters) | two homes, one updated | 153 / 42 |
Cause. Not carelessness in any single edit. Each correction was made correctly — in the place the author happened to be reading. STATE.md is ~1,350 lines, and a fact acquired a second home the moment a session summarised it in its own block while an older block still stated it. Nothing connected the two, so a correction landed in one and the other went on asserting the opposite.
Why it costs more than being silent. A file that contradicts itself cannot be partially trusted: the reader has no way to tell which half is current, so they re-derive the work anyway — which is the single thing the handoff exists to prevent. The 543a025 line is the sharpest case. A stale deployed sha reads exactly like "there is a deploy owed", and disproving it costs a VPN session and 90 seconds. It is not a missing fact; it is a fact pointing the wrong way.
Fix. src/js/state-handoff.test.js, five assertions, each falsified:
- every repo-relative path
STATE.mdnames resolves (or is exempted with a written reason — served bundle hashes and filename PATTERNS are named as evidence, not as files, and are excluded by requiring a/); - "Migrations through NNNN" matches the highest migration on disk;
- the claimed live-proof count matches what
run-proofs.mjsregisters (minusdb-query.mjs, which is the runner, not a proof); - every spelling of each of those must agree — the check runs over all matches, not the first;
- the file may state exactly one test count.
It found a third home of the test count on its first run, in a paragraph that had also been carrying "migrations through 0166" and "21 of 22 proofs green" for a week.
Where it lives now. src/js/state-handoff.test.js; STATE.md's example paragraph now carries ⛔ do not read counts out of this paragraph, and says the live numbers live in exactly one place.
The general rule. This is the repo's class 6 — two implementations of one rule drift — with prose as the implementation, and it is easier to walk into than the code version, because a document has no compiler and every sentence looks equally authoritative. Two habits, in order of value:
- ⛔ When you correct a claim, grep the whole file for its other homes before you commit. Every one of the five was a single grep away.
- Give a decaying fact exactly one home, and make the other places point at it rather than repeat it. A number in a narrative block is a copy that will never be updated, because nobody re-reads a session summary to check its arithmetic. The durable half of an old block is the LESSON; the counts in it are already wrong.
And the reason a test is the third fix here: this had been paid for at least three times before — "this line said 543a025 for a day", "two sessions repeated the locked-out claim", "it said measurement was OFF after it had been turned back on" — each time diagnosed correctly, written down, and repeated anyway.
which pg_dump said it was not installed, and it had been installed all along
Symptom (as recorded in the plan). "pg_dump, psql and the supabase CLI are all absent from this machine — which finds none of them." On the strength of that line, three phases of docs/TEAM-WORKFLOW.md were marked blocked on "install a PostgreSQL 17 client", and a note was written telling the next session to brew install libpq before planning around it.
Cause. which searches PATH. Homebrew's libpq is keg-only — it is installed under /opt/homebrew/opt/libpq/bin and deliberately NOT linked into PATH, because its binaries collide with the ones a full postgresql formula would install. So which was answering a question about PATH and the answer was read as a statement about the disk. Measured 2026-08-27: brew list --versions libpq → libpq 18.4, and /opt/homebrew/opt/libpq/bin/pg_dump --version → pg_dump (PostgreSQL) 18.4. It had been there the whole time.
There was a second, smaller error stacked on the first: the plan reasoned that a client older than the 17.6 server would refuse, and concluded "install 17". The refusal is one-directional — pg_dump refuses to dump a server newer than itself, and a newer client dumping an older server is the supported case. 18.4 against 17.6 was always fine.
Fix. Use the full path, or export PATH="/opt/homebrew/opt/libpq/bin:$PATH". docs/TEAM-WORKFLOW.md §7.1 now records the correction, and phase 1 is blocked on the database password alone.
Where it lives now. docs/TEAM-WORKFLOW.md §7.1.
The general rule. Class 7 — check that the INSTRUMENT can see the thing. A negative result is a statement about what the instrument searched, never about what exists. which searches PATH; grep searches the file you gave it (not the shared chunk the code actually landed in); a minified bundle has no module-scope names to find. Before believing a zero, ask what the tool looked at, and check it against a subject you KNOW is there — here, one brew list --versions would have cost five seconds and saved a day of the plan being wrong about its own blockers.
A pg_dump restore made the copy MORE permissive than the original
Symptom. samo-dev was built from a pg_dump of production and compared against it object by object. Tables, functions, triggers and RLS-enabled tables matched exactly. Grants did not: dev had 134 that production does not have, and 0 that it was missing. Sixteen tables — students, people, student_change_requests, student_import_batches, _timeline_backup_0166, schema_migrations and ten more — had been granted to anon, which production grants nothing on.
Cause. Supabase sets ALTER DEFAULT PRIVILEGES granting everything to anon / authenticated / service_role on newly created tables. pg_dump writes the GRANTs a database HAS; it writes no REVOKEs, because it assumes stock PostgreSQL defaults where nothing is granted to anyone. Every table the restore CREATES therefore picks up the platform's defaults first, and nothing in the dump takes them back off. The dump is a description of what is granted, not of what is denied, and on a platform with non-standard defaults those are different things.
Why it was not cosmetic. RLS was still enabled on all of them, so rows were still filtered — which is exactly why it would have survived a casual look. But production refuses anon at the GRANT, before any policy runs, and dev would have refused only at the policy. Two databases, two different gates, and docs/TEAM-WORKFLOW.md D2 puts no door gate on the preview URL precisely because dev is supposed to behave identically (§7.3).
Fix. Generate the REVOKEs from the measured difference — never from a hand-written list of tables that look sensitive — and re-measure until extra = 0 and missing = 0. It took 134 revokes. npm run dev:check (tools/dev-check.mjs) is now the ratchet: it compares the anon key's HTTP status on both databases across allow-subjects AND deny-subjects, and was falsified by granting anon SELECT on students on dev alone — it reported DRIFT and exited 1.
Where it lives now. skills/build-the-dev-database.md §3b, tools/dev-check.mjs.
The general rule. Comparing two systems, compare what each DENIES, not only what each allows. A copy is verified by the differences being zero in both directions — "everything the original had is present" is half a check, and it is the half that cannot see an addition. And when a platform ships non-standard defaults, any tool that emits only the positive state (a dump, an export, a seed) will silently inherit them.
A refresh script printed "identical to production" while refreshing nothing
Symptom. tools/dev-refresh.mjs, run for the first time, finished with row-count diffs: 0 · grants extra: 0 · grants missing: 0 and ✓ samo-dev rebuilt and identical to production. Buried above it, filtered out of the summary, was ERROR: duplicate key value violates unique constraint "users_pkey" … CONTEXT: COPY users, line 1.
Cause, two halves that hid each other. Step 2 drops and recreates public and passport — it does not touch auth, which is a different schema. So COPY auth.users hit rows from the previous load and aborted the entire COPY at line 1, leaving dev's accounts exactly as they were. Step 6 then compared 64 tables — public and passport only — and never looked at auth, so it could not see what step 4 had failed to do. The verification's blind spot was in the same place as the bug, which is why the run went green.
It only looked correct because the stale auth copy happened to be identical to the fresh one: the hand-run had loaded it minutes earlier. A refresh a week later would have carried a week-old set of accounts and still reported parity.
Fix. truncate auth.users cascade before loading, and include auth.users / auth.identities in the comparison — 64 tables became 66, which is the number the by-hand check had used all along. Re-run: 66 compared, 0 diffs, and the sign-in proof re-run against the rebuilt copy because truncating auth had wiped the password the earlier proof set.
Where it lives now. tools/dev-refresh.mjs steps 4 and 6, skills/build-the-dev-database.md.
The general rule. A verification that covers less than the operation is not a verification — and the gap is always where the bug lives, because the same blind spot produced both. Enumerate what the operation TOUCHES, then check that the verification's subject list covers all of it. Here the operation wrote to auth and public; the check read only public. Also: a script that filters its own noise must not filter its own errors. The COPY failure was on screen the whole time, three lines above a green summary that contradicted it.
npm test | grep returned success while the suite was failing
Symptom. Two commits were pushed to main on 2026-08-27 with a failing test, hours after CI was made blocking. Both were written as guarded chains that looked safe.
Cause, and the second one is the interesting half. The first was npm test ... ; git add — a semicolon, so nothing gated the commit. The "fix" committed for it was npm test 2>&1 | grep -E "Tests" && git commit — and a pipe replaces the exit status with the LAST command's. grep found the summary line and exited 0, so && saw success while the suite was red and the failure was printed on screen three lines above.
Fix. npm test > /tmp/t.log 2>&1; echo $? and read the CODE, or keep any pipe out from between the test and the &&. Direct pushes to main by an admin bypass the required check (enforce_admins: false, deliberate — it is what lets the owner push), so branch protection does not catch this shape.
Where it lives now. skills/ habits and this entry.
The general rule. Class 7 — the instrument decides what can be seen. A pipeline's exit code describes its LAST stage, not its first; a summary line matching is not the same fact as a suite passing. Same shape as which pg_dump reporting "absent" for a keg-only binary earlier the same day: both times the tool answered a narrower question than the one being asked, and the narrower answer was read as the broader one. When a check gates something, verify the CHECK reports failure — run it once against a known-bad state.
urllib got 403 from Discord and I reported the service as DOWN
Symptom. A check of the VitalSound webhooks reported all twelve DEAD with HTTP 403, and "VitalSound notifications are broken for 12 departments" was reported to the owner as a live production outage.
Cause. The check used Python's urllib.request, whose default User-Agent is Python-urllib/3.x. Discord's edge rejects it with 403 Forbidden regardless of whether the webhook exists. The same twelve URLs checked with curl — and with Node's fetch, which is what tools/discord-webhook-identify.mjs uses — returned 200 with the channel id, all twelve alive.
How it was caught. The twelve had been created and read back through the bot API minutes earlier, so "all twelve dead" contradicted a known-good observation from the same session. A result that contradicts a breadcrumb you already have is the instrument, not the world.
The earlier reading was probably wrong too. The check that first declared the OLD VS webhook dead used the identical urllib script and got the identical 403. The webhook it condemned may have been healthy; it has since been replaced, so that can no longer be settled — which is itself the cost.
Fix. Use the committed tool (npm run webhook:id), which uses fetch. Never hand-roll an HTTP check against a third-party API with a library whose default User-Agent is a bot signature.
Where it lives now. tools/discord-webhook-identify.mjs.
The general rule. Distinguish "the service says no" from "the service did not answer the question you think you asked." 401 and 403 are different answers: 401 said "invalid webhook token" (a real verdict about the credential), 403 said nothing about the credential at all. Before believing a negative result from a network probe, reproduce it with a second client. Two clients disagreeing means the instrument; two agreeing means the world.
The verification command in STATE.md named a sha two deploys behind — so following the handoff's own instructions reported a deploy that had already shipped
Symptom. STATE.md said, in bold, "Check, do not trust this line", and gave the command to run. Running it printed seven changed files and 132 insertions under src/ — the unmistakable shape of a deploy is owed. Nothing was owed. Everything it listed had shipped two deploys earlier.
Cause. The deployed sha had four homes in one file, and exactly one of them had been corrected:
| line | said | actual |
|---|---|---|
| the ✅ DEPLOYED line | 2151d6a | ✅ correct |
| "Previous:" | 36ac1d5 | 832bb14 (two deploys stale) |
the git diff snippet, twice | 7405712 | two deploys stale |
| the closing "no deploy is owed" | 7405712 | two deploys stale |
This is class 6 (two implementations of one rule drift) with a twist that made it much more expensive: the stale copy was the INSTRUMENT. The file did not merely assert something false — it handed the reader a working command that produced convincing false evidence, complete with a diffstat. A prose claim invites doubt; a command's output does not.
state-handoff.test.js had a comment in its own header naming this exact failure ("the deployed sha named a commit two deploys behind... costs a VPN session to disprove") and had not been given an assertion for it. A hazard written down in the guard's comments is not guarded.
Fix — remove the retyping, do not retype more carefully. npm run deploy:owed (tools/deploy-owed.mjs) parses the ✅ DEPLOYED line — the sha's one home — and diffs it against the working tree. The guard then forbids the shape: STATE.md may not contain git diff <sha>..HEAD at all, and must declare exactly one DEPLOYED sha, which must resolve to a real commit.
A second bug, found by falsifying the first. deploy-owed.mjs v1 used git diff <sha>..HEAD, copied from the snippet it replaced. That compares commits, so with an edited src/main.css sitting unstaged it answered "NO DEPLOY OWED" — the instrument could not see the hazard, and an uncommitted shipping change is the more urgent kind, since it is not even pushed. Omitting ..HEAD diffs the deployed commit against the working tree; git ls-files --others catches a file never added at all. Only the ritual found this: reintroduce the bug, watch it fail, restore.
A third, in the guard's own exemption list.ABSENT_ON_PURPOSE['src/html/tab-golden-period.html'] said "PLANNED, not written — DELETE this exemption in the same commit that creates the file." The file was created; the exemption stayed. For every day after, the dead-pointer sweep skipped a path that existed — rename or delete that file and both sweeps stay green while STATE.md points at nothing. Now asserted: no exemption may survive its file arriving.
Where it lives now. tools/deploy-owed.mjs · npm run deploy:owed · three new assertions in src/js/state-handoff.test.js.
The general rule. When a fact is retyped into a command, the command is a copy that rots — and a rotten instrument is worse than a rotten sentence, because its output looks like evidence. Delete the copy: have the command READ the fact from its one home, and let the guard forbid the shape that reintroduces it. And an exemption is a claim about the world too — the ones that say "not written yet" expire, and a guard whose exemption outlives the absence it describes fails GREEN.
"The VM can't do mail" — one probe answered a narrower question than the sentence it was written into
Symptom. An assessment concluded do not self-host email, and the owner immediately asked the obvious question back: "isn't there a way to send email from the VM?" There is. smtp.gmail.com:587 answers from that box, completes STARTTLS, and offers AUTH. The conclusion had been generalised past its evidence.
Cause. One probe was run — can something outside connect IN to the VM? — and the answer, correctly no, was written up as though it settled three different questions:
| question | direction | truth |
|---|---|---|
| can anything connect in? | inbound | no — only 443 is mapped |
| can it deliver direct to MX? | outbound :25 | no — egress blocked |
| can it send via a relay? | outbound :587 | YES, and never tested |
The probe swept 202.28.95.46 — the public address. Nothing about it could have answered a question about egress, because it was pointed the wrong way. The write-up then reached a conclusion that needed all three.
The tell was in the evidence and went unread: curl https://github.com had already returned 200 from that box in the same session. Outbound worked, it was recorded, and it was not connected to the claim being made.
A second error inside the correction. The follow-up probe reported smtp-brevo.com:587 blocked. That host does not exist — Brevo's is smtp-relay.brevo.com, which is OPEN. A DNS failure and a filtered port are different facts and the probe printed them identically, so an invented hostname became a finding. Resolve the name first and print the IP; a probe that cannot distinguish "no such host" from "blocked" will manufacture blockers.
Fix. Test each direction separately and say which one each result belongs to. docs/EMAIL.md §3 is now a three-row table — send / be a server / receive — because those were always three answers wearing one sentence. And a TCP connect is not a service: the relay claim is backed by an actual SMTP session (openssl s_client -starttls smtp) showing the 250-AUTH line, since a captive proxy will complete a handshake and nothing else.
Where it lives now. docs/EMAIL.md §3.
The general rule. A probe answers the question its direction asks, not the sentence you write around it. Before generalising a negative result, name the question it actually tested and check whether the conclusion needs a wider one — "X cannot do Y" almost always hides an unstated direction, endpoint or credential. Two supporting habits, both paid for here: evidence already collected in the session (that github 200) is evidence against your claim too, so re-read it; and resolve a hostname before reporting its port shut, because a typo and a firewall look identical from a connect() call.
A dashboard was about to report 83% of a quota that was really at 7% — sentinel strings, and a data import counted as traffic
Symptom. A new สถิติ panel was ready to show 25 Apps Script calls in one minute against a 30-simultaneous ceiling. The owner asked the only question that mattered: "are there really 25 pr upload in 60 seconds?" There were not. The real peak is 2.
Cause 1 — a sentinel is not a value. The count used file_url is not null as "has an upload". That column holds four different things:
| value | rows | is it an upload? |
|---|---|---|
https://drive.google.com/file/d/… | 98 | yes |
null | 61 | no |
ลิงก์เสริม: <url> — a link the submitter PASTED | 50 | no |
ไม่มีไฟล์แนบ — "no attachment" | 9 | no |
null was handled. The two Thai sentinels were not, so 98 real uploads were reported as 157. This is the repo's own recurring shape wearing new clothes: asking whether a field is null instead of whether it resolves.
Cause 2 — a timestamp records when a ROW was written, not when work was done. The 25 rows landed within 2.86 seconds, ~65 ms apart. That is the Sheets→Supabase migration writing straight to Postgres, for files that were already in Drive. Not one Apps Script call happened. Nothing in a count() distinguishes a bulk import from a stampede — the shape is identical.
Fix. Count only real uploaded files, and drop rows arriving under a second after the previous one: no human submits two forms 65 ms apart, and that spacing is the signature of a machine. excluded_bulk ships in the payload so the exclusion is visible — an exclusion nobody can see is how a number quietly becomes a lie.
How close this came to being expensive. 83% of a ceiling is a number somebody acts on. The panel would have argued for serialising uploads, splitting the Google account, or a migration off Apps Script — real work, to fix a system sitting at 7%.
A near-miss in the same session. A build slip left the corrected migration containing only comments and one comment on function statement. apply-migration.mjs printed ✓ migration applied, because that statement succeeded. The function was untouched and the old numbers kept coming back. Every apply is now followed by reading the live body back (pg_get_functiondef(oid) like '%<new marker>%').
Where it lives now. supabase/migrations/0173_gas_count_real_uploads_not_sentinels_or_imports.sql · src/js/analytics-email.test.js.
The general rule. A derived metric is a claim about the world, and it must be checked against the ROWS before anyone is shown it. Both errors survived being written, reviewed, and a passing test suite — and neither survived thirty seconds of select … limit 26. Before shipping an aggregate, print the records behind its most extreme value and look at them. Two questions catch this class: what else can this column contain? and what would a bulk write look like here?
Impersonating a user through the Management API works for one statement and silently stops working at the next
Symptom. Testing analytics_overview() needs a caller with an admin grant — as the Management-API superuser auth.uid() is null, so the RPC correctly raises. The documented workaround works:
sql
begin;
select set_config('role','authenticated',true),
set_config('request.jwt.claims', json_build_object('sub', '<uuid>')::text, true);
select public.analytics_overview(90); -- ✅ returns data
rollback;The identical pattern with an update in place of the select silently does nothing. No error, HTTP 200, and the row is unchanged.
Cause. The endpoint does not carry the transaction-local settings across statements the way a psql session does. Verified rather than guessed: a probe inside the same block returns role = authenticated, is_staff = true, auth.uid() non-null — the impersonation genuinely takes effect — it just does not survive to the next statement, so the update runs back as superuser and users_self_update_guard rejects it.
What made it hard to see. The failure has two layers of camouflage. The API returns only one result set, so a multi-statement block can answer with the set_config row and look like a success; and public.users has been SELECT self-only since 0147, so returning email yields nothing even when an update does land. "No rows back" therefore means either refused or invisible.
Fix. For READS of a guarded RPC, the pattern is fine — one statement after the config, inside one block. For WRITES from a maintenance script, use set session_replication_role = 'replica' (what dev-refresh.mjs already does to load data) and make the script refuse every project except the disposable one, by ref, before it writes — see tools/dev-grants.mjs.
Where it lives now. tools/dev-grants.mjs (the reasoning sits beside the line it explains) · src/js/dev-grants.test.js asserts the refusal ordering.
The general rule. A session setting is not a session when the transport re-connects between statements. Before trusting an impersonated write, read the row back as superuser in a separate call — the write path's own answer cannot distinguish "refused" from "invisible to me". And treat a multi-statement block against an HTTP query endpoint as one statement's worth of guarantees.
npm run proofs against dev ran two proofs against PRODUCTION and printed one green summary
Symptom. Nobody reported it — it was found while building the CI job that would have relied on it. The documented way to point the proofs at the dev database,
bash
VITE_SUPABASE_URL=$SUPABASE_DEV_URL npm run proofssent the 17 .sql proofs to samo-dev and proj0092-seat-parity.mjs + grant0093-reads.mjs to production, then printed all 25 proofs green over the mixture. Nothing in the output named a project.
Cause. 39 tools in tools/ each hand-rolled the same .env.local parse, and they did not agree. db-query.mjs had ALREADY been bitten by this on 2026-08-28 — a migration applied to dev, verified with the tool, reported NOT APPLIED, because the check read a different database than the write — and it was fixed there by letting process.env win. Its siblings were not fixed, and the header recording the lesson sat in the one file that no longer had the bug. A fix in one file is not a mechanism.
The second half of the same defect: .env.local was read with an unguarded readFileSync, so any environment without that file (a CI runner) got an unhandled ENOENT that reads like a broken tool rather than a missing credential. Measured: 21 of 23 proofs failed a CI-shaped run that was holding perfectly valid credentials in its environment.
Fix. tools/env-lib.mjs — one loader (.env.local optional, process.env wins) and one resolveTarget() that derives the label by comparing refs, never by trusting which variable a value came from. Adopted by db-query.mjs (which covers every .sql proof and the four .mjs proofs that shell to it) and by the two that hand-rolled HTTP.
But the mechanism is not the refactor. run-proofs.mjs now reads each proof's own → project: <ref> announcement back and FAILS any proof whose answer came from a database other than the one it was sent to; a proof that announces nothing is UNKNOWN, never PASS. That catches a proof nobody has written yet, which the two-file fix does not. --dev additionally refuses to run if SUPABASE_DEV_URL does not resolve to samo-dev, and SKIPS the two non-database proofs explicitly with the reason printed — a summary that silently shrinks is its own bug.
Where it lives now. tools/env-lib.mjs, tools/run-proofs.mjs, .github/workflows/proofs.yml, guarded by src/js/proof-targeting.test.js. Both runtime branches were falsified by reintroducing the drift (FAIL: "ran against fheueuowbchsnsvbcgil, not the samo-dev ref") and by removing the announcement (UNKNOWN), and the static guard by making a proof parse .env.local again.
⚠️ The static guard's first version fired on the healthy case — it required every .mjs proof to reach the database through env-lib, including repo-protection.mjs and notify-exposure.mjs, which ask GitHub and Cloudflare and hold no database credential at all. Its subject is now derived from the runner's own NON_DB set.
The general rule. A tool that can be pointed at more than one database must SAY which one answered, and whatever aggregates those tools must check the answer came from where it was sent. Writing the lesson into the header of the one file you just fixed leaves every sibling holding the bug — and when the aggregate prints a single verdict, the mixture is invisible by construction. Corollary paid for here twice over: before trusting a green suite, ask what it would look like if half of it had answered from somewhere else.
main's CI was red for a day because a guard could not see the commit it was checking
Symptom. Every build run on main had failed since 2026-08-28 with
Error: STATE.md says DEPLOYED = e0bd2e2, which is not a commit in this repo.and then, a day later, the same sentence with f9584e5. Both shas were perfectly correct and present. Nobody noticed, because a check that is always red is indistinguishable from a check.
Cause. actions/checkout@v4 fetches depth 1. state-handoff.test.js verifies STATE.md's deployed sha with git cat-file -e <sha>^{commit}, and in a shallow clone every commit but the tip is simply absent — so git answers exactly what it answers for a MISTYPED sha. The guard already handled "no git at all" (a tarball, a sandbox) and returned inconclusive; it had no idea that a git which is present can still be unable to see a valid object.
Why it mattered more than a red X. build is a REQUIRED status check on main (phase 0's highest-value guardrail, enabled 2026-08-27). A permanently false red there blocks every contributor PR — the guardrail built to protect the branch was quietly closing it.
Fix. Both halves, because either alone is a half-fix: build.yml checks out with fetch-depth: 0, so the guard is REAL in CI rather than merely quiet; and the test now asks git rev-parse --is-shallow-repository before failing, so "I cannot see that object" is never reported as "that object does not exist". Falsified both ways — a bogus sha in a full clone still fails, and the correct sha in a real --depth 1 clone passes.
Where it lives now. .github/workflows/build.yml, src/js/state-handoff.test.js.
The general rule. A guard's instrument needs a guard too, and the question "did it answer NO, or could it not see?" is the one that separates them. This repo already had that rule for which not finding pg_dump (a PATH answer read as a disk answer); it recurred verbatim in git. And check the CI dashboard: a guard nobody looks at fails silently no matter how loudly it prints, and this one had been shouting for a day into a tab nobody opened.
A CI gate whose red depended on jsDelivr, and two tests that passed over deleted code
Symptom. None yet — all three were found by a scrutiny pass on the day they were written, before anyone was misled. They are recorded because each is a recurrence of a rule this repo already had.
1. The browser smoke could go red on a CDN blip. smoke-browser.mjs failed on ANY failed Script/Stylesheet fetch, and the page loads from cdn.jsdelivr.net, cdn.quilljs.com and fonts.googleapis.com. A slow route from a GitHub runner would have turned the check red for a reason unrelated to the change — a warning that fires on the healthy case, reproduced inside the instrument built to enforce that very rule. Now only our own origin fails the run; a third-party failure prints a note.
⚠️ This is safe only because a real outage is still caught, by a DIFFERENT check. Blocking jsDelivr was measured: the page breaks badly (Bootstrap is a classic CDN script, so the module graph throws and no handler binds) and the BOOT checks fail. Measure the outcome, not the cause — the network assertion added little and carried all the false-positive risk.
2. The same file's watchdog diagnostic named the wrong culprit. It printed "the watchdog fired on a page that booted — it is crying wolf again" over a page that had NOT booted, where the watchdog was right. The check is now skipped, and says so, when the page did not boot.
3. Two guards passed while the code they guard was deleted. Both were vacuity, both in tests written that same hour: · one asserted an identifier appeared in auth.js — and it still appeared, in a comment. Fixed with stripComments(), which four other tests already use. · one asserted branch ORDER with indexOf(a) < indexOf(b) — and when a was deleted, indexOf returned −1, which is less than everything. Fixed by requiring both to exist first.
4. The proof-target check read only the FIRST announcement. A proof querying dev and then production would have passed on its opening line. Now every announcement must match; falsified by making a proof emit a second one.
Where it lives now. tools/smoke-browser.mjs, tools/run-proofs.mjs, src/js/google-provider-guard.test.js, src/js/proof-targeting.test.js.
The general rule. A guard written in the same hour as the code it guards has not been tested against anything but the author's intent — delete the code and watch it go red, and treat every indexOf, substring or identifier match as vacuous until you have seen it fail. And when a check can be red for a reason outside the change, prefer the check that measures the OUTCOME: the boot probe catches a CDN outage without ever asserting on the network.
npm run deploy:owed said "production is serving current code" while /docs was a whole rebuild behind
Symptom. Right after the docs site was restructured and pushed, npm run deploy:owed printed ✅ NO DEPLOY OWED — production is serving current code. It was wrong: samo.md.kku.ac.th/docs/start/install answered 404 on the VM while main had contained that page for an hour.
Cause. SHIPPED in tools/deploy-owed.mjs was ['src/', ':!src/**/*.test.js', 'index.html', 'admin/index.html']. That was a complete list on the day it was written — a deploy only published two bundles. Earlier the SAME DAY, server/deploy.sh learned to build the docs with DOCS_BASE=/docs/ and publish them to /var/www/docs, and nobody widened the list. The one instrument that answers is production current had gone blind to half of what production serves, and it answered GREEN, which is the direction that gets believed and acted on.
Fix. docs/ (minus docs/state/** and docs/state-archive/**, which are notes rather than published artifacts and would make it cry wolf on every handoff), plus server/nginx-samo.conf and server/deploy.sh — a config change also needs a trip to the VM. The reason is written beside the list.
Where it lives now. tools/deploy-owed.mjs, in the comment above SHIPPED.
The general rule. When the deploy learns to publish something new, add it to the deploy checker in the SAME COMMIT. An instrument's subject list is correct only for the system it was written against, and nothing tells you when the system grows past it — the checker keeps answering confidently about the part it still knows. This is the sibling of "a guard cannot see the hazard" (§write-a-guard trap #1): here the guard could see fine, it had simply never been told the building had another floor.
Dead-link checking on the docs site had been off since the day it was built
Symptom. None — found while sweeping for damage after a large restructure, before anyone was misled. The docs site built green through a change that moved or deleted every contributor-facing page.
Cause. ignoreDeadLinks in the VitePress config carried /^\/(?!samomdkkuweb)/, intended to mean "ignore absolute links that are not part of this site". It does the opposite of what its author believed. In markdown you write a site link WITHOUT the base — /contributing, /start/install — and VitePress prepends base at build time. So no link in any source file ever begins with /samomdkkuweb, the negative lookahead matched every one of them, and the whole site's internal links were exempt. A restructure could have shipped with completely broken navigation and a passing build.
Fix. The pattern is deleted. Verified in both directions: the build still passes with it gone (so there were no actual dead links), and a deliberately inserted [x](/no-such-page-xyz) is now reported and fails the build.
Where it lives now. docs/.vitepress/config.mjs, with a ⛔ note so it is not reintroduced by someone reasoning about it the same way.
The general rule. An ignore-pattern is a guard running in reverse, and it is never tested. A wrong assertion goes red; a wrong exemption goes quiet, and its blast radius is everything it silently covers. When you write one, prove it with a control — insert the thing it should still catch and watch the build fail. And be specific about which STRING the pattern sees: a build tool rewrites paths, so the value in the source file is often not the value you are picturing.
git push to main refused with GH013 right after the org transfer, while the protection proof was all green
Symptom. samomdkkuweb was transferred to the samomdkku organisation. The first push of the follow-up commit was rejected:
remote: error: GH013: Repository rule violations found for refs/heads/main.
remote: - Required status check "build" is expected.
remote: - Changes must be made through a pull request.node tools/repo-protection.mjs reported six of six checks passing, including enforce_admins stays OFF — the very check whose stated purpose is "it is what lets the owner push main". The proof and the remote disagreed, and the proof was the one being trusted.
Cause — there are TWO enforcement paths, and the proof read one. GitHub enforces branch rules through the classic repos/{r}/branches/main/protection API and, independently, through rulesets (repos/{r}/rulesets). This repo has carried an active ruleset named main-protect since 2026-05-23. repo-protection.mjs never mentioned rulesets, so for three months its six green checks described half the gate, and nothing said so.
The transfer then emptied that ruleset's bypass_actors. Read back from the ruleset's own version history — GitHub keeps one, which is what turned a theory into a fact — the pre-transfer version held four:
RepositoryRole 5 (admin) bypass_mode always ← the owner's direct push
Integration 85455 / 946600 / 1143301 ← three GitHub Appsand the post-transfer version held []. Nothing reported this. The classic enforce_admins flag was genuinely untouched, so the check watching it stayed truthfully green while the capability it exists to protect was gone.
Fix. RepositoryRole 5 restored with bypass_mode: always; the push then succeeded (GitHub still prints the bypassed rule as a remote: line — that is a log of the bypass, not a rejection, and it is easy to misread as failure). The three Integration bypasses were deliberately not restored: they cannot be resolved to a named app through any public endpoint, and gh api orgs/samomdkku/installations returns total_count: 0, so they would grant nothing to nobody. Re-granting an unidentifiable app the right to bypass main is not a thing to do to make a red line go green.
repo-protection.mjs now reads both paths, and asserts the bypass by property — "an admin can still push" — rather than re-listing whatever is configured today. It also asserts a ruleset governs main before asserting anything about its bypasses, so deleting the ruleset fails loudly instead of passing over an empty list.
Proved by the ritual: bypass stripped → ✗ admin can still push main (ruleset bypass): expected true, got false → restored → green.
Where it lives now. tools/repo-protection.mjs, with the cost written above the code. skills/move-the-repo-to-an-organisation.md §5d.
The general rule. Ask whether the thing you are reading is the thing that DECIDES. A proof named after a capability ("protection") but bound to one API is a proof about that API, and it will keep answering confidently after the platform grows a second mechanism beside it. This is the class-7 failure in its most convincing form: not a broken assertion, but six true ones that together imply something false. When a proof and the real system disagree, the proof is the suspect — and when a platform offers two ways to configure one behaviour, a guard that reads one of them is guarding nothing.
Corollary, for any transfer or migration: capability is not configuration. The classic flag survived and the capability did not. Test what you can still DO, not what the settings still SAY.
deploy:owed went blind to a page production serves — again, one directory deeper
Symptom. npm run deploy:owed printed "no deploy owed" while samo.md.kku.ac.th/docs/state/phuriphatma was serving a page several commits stale. The tool exists specifically to answer "is production current", and it answered wrong in the direction nobody investigates: green.
Cause. Two spellings of "what production serves", drifted.
| says what ships | |
|---|---|
docs/.vitepress/config.mjs | collect() globs every .md under docs/; srcExclude names only node_modules, templates, package, demos |
tools/deploy-owed.mjs | SHIPPED carried ':!docs/state/**', ':!docs/state-archive/**' |
The exclusion was reasonable-sounding — a person's session notes are not "shipping" in any sense the author cared about — and simply untrue: VitePress globs them, the sidebar links them, nginx serves them. Verified with curl, not by reading either file.
This is the SECOND time in two days. On 2026-08-31 the same tool watched only src/ and answered "production is current" while /docs was an entire restructure behind. That fix added docs/ and wrote a header saying "WHEN THE DEPLOY LEARNS TO PUBLISH SOMETHING NEW, ADD IT HERE IN THE SAME COMMIT" — a paragraph, in the file, which did not stop the same class recurring inside the directory it had just added.
Fix. The exclusions are gone. And because a header saying "keep these in step" demonstrably does not keep anything in step, src/js/docs-shipping-parity.test.js now asserts the PROPERTY at npm test time: every page VitePress actually publishes is a page deploy-owed can see change. It calls collect() — the real publisher — rather than re-reading the config's exclusion list, so it cannot be satisfied by two lists that are wrong in the same way. Two controls guard it: docs/ must be watched at all, and collect() must return a non-empty set. Reintroduced the exclusion, watched it name state/phuriphatma.md, restored.
If those notes should not be public, the fix is srcExclude — and then they may leave SHIPPED. Either arrangement passes; disagreement does not.
A second, smaller trap from the same hour. Verifying the by-hand docs publish, grep -c prof_can_see_document on the served page returned 1 and I read that as "the new section shipped". It had not: the phrase occurs in an OLDER section of the same page. The control (Two wrong turns) was also 1, so both halves agreed — on the wrong answer. A verification string must be unique to the change, not merely present in it; re-grepping for The sweep afterwards returned 0 and told the truth.
The general rule. When a deploy learns to publish a new path, the tool that reports staleness must learn it in the same commit — and the only durable way to make that happen is a test that asks the PUBLISHER what it publishes. Any instrument whose subject is a hand-maintained list of paths will go blind the first time somebody adds a path; derive the list from the thing that actually does the work.
A proof that was GREEN by hand and RED under its own runner
Symptom. tools/dept0179-kinds.sql passed 10/10 when run directly through node tools/db-query.mjs. Registered in run-proofs.mjs and run by npm run proofs, the same file against the same database reported:
dept0179-kinds.sql ✗ FAIL FAIL in blobCause — the runner reads the LAST result set, and mine was a summary.run-proofs.mjs looks for a per-case verdict column. My script ended with
sql
select count(*) filter (where expected = got) as passed, ...which has no verdict column, so the runner fell through to its blob branch: if (/FAIL|DENIED_UNEXPECTED/i.test(text)) return FAIL. That branch scans the raw output text — and found the word FAIL inside the file's own CASE expression, case when expected = got then 'PASS' else '*** FAIL ***' end.
So the proof was failed by a string it printed while describing how it would report a failure. Nothing was wrong with the assertions or the database.
Fix. Shaped like dept0177-page-scope.sql, which was already correct: one row per case, a result column, and nothing after it.
sql
select case_name as step,
case when expected = got then 'PASS' else 'FAIL' end as result,
expected, got
from probe order by case_name;Where it lives now. tools/dept0179-kinds.sql, with the reason in a comment above the final select so the next person does not "improve" it by appending a summary again.
The general rule. A proof is only as good as the runner's ability to READ it, and the runner's contract is a shape, not a sentence. Two things generalise beyond SQL. First, a trailing summary row hides the cases — anything that aggregates after the per-case output moves the answer out of the shape the harness parses. Second, and worse, a fallback that scans raw text for a word will find that word in the code that produces it; a heuristic branch cannot distinguish a verdict from a string literal describing verdicts. If you write such a fallback, it must be the last resort and it must say so loudly — which this one does, and the fix is on the proof's side, not the runner's.
⚠️ The tell is the one that matters: green by hand, red under the runner, same subject. That difference is never the database. It is the instrument.
A check whose own audience could not run it
Symptom. docs/start/install.md told a brand-new contributor to verify their setup with:
bash
npm run dev:checkRun by the person it was written for, that exits 1 with
✗ PRODUCTION: URL or anon key missing from .env.localhaving proved nothing about the four keys they had just pasted.
Cause. tools/dev-check.mjs:51 reads VITE_SUPABASE_URL and VITE_SUPABASE_ANON_KEY — it is a PARITY guard, comparing samo-dev against production, and needs both sides by design. A contributor has neither production value and must never be sent them (.claude/rules/security.md).
The command was correct. The audience was wrong, and nothing in either the script or the guide connected the two.
Why this is worse than a missing check. A verification step that fails on a CORRECT setup blames the reader for the guide's mistake, at the exact moment they have no way to tell which of the two is at fault. A newcomer's first five minutes is the worst possible place to spend that confusion.
Fix. A second, contributor-facing command — npm run env:check (tools/env-check.mjs) — which asks only what its reader can answer: the file exists, the four SUPABASE_DEV_* names are present, none is still the placeholder from .env.local.example, no value got wrapped in quotes, and the dev database answers. Each failure says what to do about it.
⛔ The two were deliberately NOT merged. Teaching dev:check to skip the production half when credentials are absent would make it pass in the one case it exists to catch — a guard that fails GREEN. env-example.test.js pins both directions: env-check reads no production name, dev-check still reads two.
Where it lives now. tools/env-check.mjs, docs/start/install.md §4d (with a warning naming the other command), README.md's script table, and src/js/env-example.test.js.
The general rule. Before telling someone to run a command, check it works with the credentials THEY have. A tool's requirements are part of its contract and are invisible in its name; dev:check sounds like it checks your dev setup. The wider shape: a diagnostic must be runnable by the person the diagnosis is for, or it converts a solvable problem into a mysterious one.
A guard that asserted the implementation, and went red on a refactor
Symptom. dept-content.test.js failed with "a row created from หน้าฝ่าย is public the moment it is added" — a real and serious-sounding message — while the property it names was perfectly intact.
Cause. The assertion counted string literals:
js
const kinds = insert.match(/visible:\s*false/g) || [];
expect(kinds.length).toBe(2); // one per kindIt was written when addRow had two branches, one per kind. 0179 added two more kinds and folded all four seeds into one insert with a single visible: false, so the count went to 1. Nothing regressed; the guard was pinned to the SHAPE of the code rather than to what the code must be true of.
Fix. Assert the property, plus the one way it could be undone:
js
expect(insert).toMatch(/visible:\s*false/);
// a seed spread AFTER visible:false could override it back
expect(insert).toMatch(/visible:\s*false,\s*\.\.\.seed/);Where it lives now. src/js/dept-content.test.js, "every row addRow creates is a draft, whatever the kind".
The general rule. A guard written against today's shape fails on a refactor and says nothing about a real regression — and it costs more than the false alarm, because the fastest way to make it green again is to change the number, which is how a guard quietly stops meaning anything. Ask what must be TRUE, not how many times it is currently written down. Related and already here: never write a guard from the same list the code came from.
The browser smoke covered the smaller entry, and nobody noticed for months
Symptom. None — that is the point. Found on 2026-09-02 by asking a question nothing in the repo asks: after adding a static import to dept-page-admin.js, did anyone actually LOAD the admin page? The answer was no, and there was no way to have done it, because tools/smoke-browser.mjs only ever loaded /.
The asymmetry. This app has two entries. The public one is 226 KB. The admin one is 537 KB — most of the code in the project — and it had zero browser coverage. Everything that looked like coverage stops short of the thing that breaks:
| What it proves | What it cannot see |
|---|---|
npm test (jsdom, no bundle) | a module-level throw in a built entry |
npm run build | that it COMPILES, not that it RUNS |
curl | grep <string> | a string is present in a file that never executed |
So a dept-visual-editor.js with a top-level error would have passed 1,712 tests, a clean build, a served-artefact grep and a 13/13 smoke — and left every admin page dead. This repo has already shipped that exact failure once: the entry module never ran, Bootstrap's CDN kept every menu opening, and ~90 inline onclick="global()" handlers were silently gone.
Fix. Three checks appended to smoke-browser.mjs: the admin entry's __samoBooted, that the sign-in gate rendered, and no horizontal overflow. Deliberately run signed out — the entry module must parse and run before the gate can paint, so a boot failure is visible with no credential, which keeps the smoke safe for CI on a public repo. .github/workflows/smoke.yml already runs this file on every preview, so the new checks came with a trigger attached.
Falsified before being kept: asking for a flag that is never set turns it red with the right message (15 passed, 1 failed); restored, 16/16.
Where it lives now. tools/smoke-browser.mjs, the block after the tool-frame checks.
The general rule. Coverage is a claim about a SURFACE, and a suite that grew up around one entry point will quietly exclude the other. Count the entries — bundles, binaries, routes, workers — and ask which ones your smoke actually visits. The tell here was that the number nobody could answer was not "is it broken" but "has anyone looked", and the honest answer to the second is what surfaced the first.
"Unset is SAFE" — a guard that would have called a preview pointed at production a PASS
Symptom. None. node tools/repo-protection.mjs printed all green, and had been printing all green, while carrying an assumption that was false of one of the three projects it checks. Found by reading the sibling app's source for an unrelated reason.
Cause. The Cloudflare check judged each project's database with:
js
// Unset is SAFE — a build with no Supabase URL reaches no database.
const url = vars.VITE_SUPABASE_URL?.value;
… !url || url === devUrl …That comment is true of the app it was written against and false of the one beside it. samoweb's src/js/db.js reads import.meta.env.VITE_SUPABASE_URL with no fallback, so unset really does reach nothing. samomdkkupassport's js/app.js reads
js
import.meta.env?.VITE_SUPABASE_URL || "https://fheueuowbchsnsvbcgil.supabase.co"— the production project, hardcoded, with the production anon key on the next line. So for that project unset means production. Delete the variable in the dashboard and the preview talks to real student data, while the guard whose entire purpose is catching that reports PASS. The repo has already had the live incident this describes (refactorsamomdkkuweb holding the production URL with VITE_ENV_NAME unset); this is the same incident with the alarm disabled.
The fallback is not itself a mistake — its comment explains it exists so a missing build env cannot send passport to the RETIRED project B and split-brain the data. The mistake is a guard in repo A reasoning about what a missing value means in repo B.
Fix. Do not special-case the sibling. Require the variable to be SET and equal to the dev URL for every project: true of all three today, needs no knowledge of any app's fallback, and a deleted variable now goes RED. The message distinguishes unset from wrongly-set, because those have different remedies.
Where it lives now. tools/repo-protection.mjs, Cloudflare section — the old comment is kept inverted, as the reason.
The general rule — class 7, "guards fail GREEN", with the tell. A guard that reasons about what a MISSING value falls back to is asserting a fact about code it cannot see. The fallback lives in another repository, another language or another team's file, and it can change without the guard noticing. Assert the POSITIVE state you require — set, and set to this — rather than enumerating the states you believe are harmless. Whenever a guard's comment says a condition is "safe" or "neutral", find the code that makes it so and check it is the only code that can. Here there were two readers of one variable and only one had been read.
A proof that was written, run by hand, committed — and never ran in the suite
Symptom. None, which is the point. npm run proofs printed "all 30 proofs green" immediately after a 31st proof was added, run against samo-dev AND production by hand, and committed. The number that is supposed to be the evidence was the thing concealing the gap: 30 was also the answer yesterday.
Cause. PROOFS in tools/run-proofs.mjs is a hardcoded list. Adding the file is not adding the line, and nothing connected the two. Same class as the mistakes index found the same day — a guard whose SUBJECT is a hardcoded name (class 7) — landing this time on the tool whose whole job is telling you which guards still hold.
Fix. The directory is the authority for what exists. An unlisted tools/*.sql and a listed-but-absent one both exit 1 naming the file.
⚠️ .sql ONLY, and the first version got this wrong. PROOFS also holds .mjs proofs, while tools/ is full of .mjs that are NOT proofs (apply-migration, db-query, …). Comparing every listed entry against the .sql listing declared all eight .mjs proofs missing. A directory can only be the authority for an extension where every file is a proof.
Then it failed for a second, unrelated reason. Once it ran, the new proof reported FAIL — 6 failed while every row it emitted said "verdict":"ok". judge() counts a case as passing only if the verdict starts with PASS. Before that it had emitted RAISE NOTICE, which the Management API does not return at all — [], every case executed and none readable.
So one proof failed three ways in a row while being correct: invisible to the runner, unreadable by the runner, then misread by the runner. The SQL was right every time.
Where it lives now. tools/run-proofs.mjs (the directory check), tools/passport0180-season-gate.sql (rows, and PASS/FAIL verdicts).
The general rule. A new guard is not wired in until you have watched the runner report it. Run the suite and find your proof's name in the output — "the suite is green" is not evidence when your proof may not be in the suite. And when a proof is green by hand and red under the runner on the same database, stop looking at the SQL: the difference is never the subject, it is the instrument.
pass-hardening reported 9 failures and the database was innocent every time
Symptom. tools/pass-hardening.mjs printed 45 passed, 9 failed against production. The failures read like a live passport breach — a student could update their own total_km, edit another student's row, delete someone else's scans, call admin_leaderboard(), and read the whole roster:
FAIL cannot set own total_km -> rows=1
FAIL cannot touch another student -> rows=1
FAIL admin_leaderboard refused -> ALLOWED
FAIL reads only own profile -> 637The first diagnosis was wrong, and it was written into the handoff. It said the failures were an artefact: the Management API connects as the Postgres superuser, which bypasses RLS, so every "cannot" is expected to be allowed. Plausible, never tested, and it wrote off nine failures for weeks.
Two measurements killed it. A control through the same API showed impersonation works perfectly — set_config('role','authenticated') plus JWT claims moves current_user to authenticated, drops rolbypassrls to false, resolves auth.uid(), and cuts a passport.profiles read from 637 rows to 1. And the script's own output already disproved it: the anon block passed 15/15, including reads 0 profiles. A global RLS bypass cannot spare anon and hit only the student.
Cause. The probe's "student" was an admin. The selection reads:
sql
-- students: real kkumail auth accounts that are not admins ...
select u.id into v_student from auth.users u
where lower(u.email) like '%@kkumail.com'
and u.id is distinct from v_admin_full and u.id is distinct from v_admin_scoped
order by u.created_at limit 1;The comment says not admins; the SQL excludes only the two admins it had already picked for the admin tests, then takes the oldest kkumail account — which is a founder. It selected an account whose permissions column is empty while managed_permissions holds master, so passport.is_admin() returned true and the "student" satisfied profiles_read_self_or_admin and profiles_update_admin. All six student failures, plus an admin one and a cascading total_km delta, fall out of that single wrong subject.
Fix. Ask the gate, never a column. The loop impersonates each candidate and keeps the first two for which passport.is_admin() is false, raising if it cannot find two. Result: 54 passed, 0 failed on the same database, same minute. The ninth failure was separate and also stale — it asserted the shared passportadmin@samomdkku.app account exists, but the 17 shared-password admins were deleted on purpose on 2026-08-17/18, so a finished security cleanup scored as a broken proof. Inverted: absence passes, and a re-provisioned shared-password admin now fails loudly.
It then joined npm run proofs, which took two more fixes it had been dodging for its whole life: it parsed .env.local itself (so it ignored --dev targeting — proof-targeting.test.js fails the build on exactly this) and it never announced its target, so the runner scored it UNKNOWN. Both now come from env-lib.mjs like every other proof.
Where it lives now. tools/pass-hardening.mjs (the is_admin() loop and the inverted shared-account check), listed in PROOFS in tools/run-proofs.mjs.
The general rule. A deny-check is only as good as the identity it denies. Before believing a DENY failure, ask what the probe's subject actually holds — derived from the gate's own predicate, not from the column you expect. Here permissions was empty and managed_permissions held master; a subject check reading permissions alone would have agreed the account was an ordinary student. And an explanation that dismisses a failure needs the same proof as one that reports it — "it's just an artefact" was accepted for weeks while the script's own anon block sat three lines above, contradicting it.
The install guide explained a permission failure that a public repo cannot have
Symptom (reported by the owner, reading the guide as a newcomer would): "is this can occur?? isn't it public repo — If it says Repository not found or permission denied".
docs/start/install.md step 2 carried a prominent warning box:
If it says
Repository not foundorpermission denied— that means you have not been invited to the project yet.
Cause. samomdkku/samomdkkuweb is public (gh repo view --json visibility → PUBLIC), so gh repo clone succeeds for anyone with a GitHub account and there is no invitation to lack. The sentence was never measured; it was written by analogy with private repositories, in the voice of a finding.
Two costs, and the second is the expensive one:
- It misdiagnoses. The failures that actually happen there are a typo in the slug and
ghnot being signed in — neither of which the box mentions, so the reader stops looking at the two things it really is. - It sends the reader down the fork road, which is real and correct but loses the automatic preview site, to solve a problem they do not have. The guide's own Prerequisites §4 already said nobody has to be invited. Two homes, one corrected — class 6.
Permission does bite in this flow; it bites one step later, at git push into the repository, which is where the fork advice belongs.
Fix. The box now states the repository is public, names the two failures that do occur, and moves the fork route to the sentence about sending a change back, linking to the side-by-side table rather than restating it.
Where it lives now. docs/start/install.md step 2.
The general rule. Prose about an ACCESS failure must name the access it tested. "Permission denied means you were not invited" reads as a finding and is really an assumption about a repository's visibility — a fact one command answers. This is the "X cannot do Y hides an unstated direction, endpoint or credential" trap in .claude/rules/mistakes.md class 7, wearing documentation's clothes: the reader cannot tell an explanation from a measurement, so they believe it and stop measuring. And a guide that explains a failure the system cannot produce is worse than one that says nothing, because it hands the reader a confident wrong answer at the exact moment they are least able to judge it.
The setup guide's four keys: two the app never read, two nobody should have had
Symptom. Reported as a plain question — "so how should i send them the credentials of .env.local" — while asking whether SOPS was the right tool. Tracing what those credentials actually did answered a different question.
Cause, part 1 — the app never read them. docs/start/install.md had a contributor fill in SUPABASE_DEV_URL / SUPABASE_DEV_ANON_KEY. src/js/db.js reads VITE_SUPABASE_URL / VITE_SUPABASE_ANON_KEY. Nothing mapped one to the other — not vite.config.js, not tools/dev-all.mjs. Measured by building the file the guide produces and asking Vite's own loader:
loadEnv('development', dir, 'VITE_') -> {}So the traced sequence was: paste the values → npm run env:check prints "✓ You are set up" → npm run dev starts an app with no database. Identical to §4e's failure state. Worse on the passport half: passport/js/app.js falls back to a hardcoded production URL and anon key when the variable is unset, so a volunteer following the guide was reading the real student database.
Why nobody saw it for weeks. A maintainer's own .env.local carries VITE_* pointing at production, so npm run dev worked on the only machine it was ever run on — by talking to the live site. Both halves of that are wrong, and the working case hid the broken one.
Cause, part 2 — the guard could not see it. env-example.test.js asserted that the example declares four names and that the docs name the same four: a list against a list, both derived from the same decision. The hazard lived in the GAP between the example and the app, which no list comparison can reach.
Cause, part 3 — the four were not four of a kind. Two are the pair the built website already publishes (an address and a visitor key RLS gates). The other two are SUPABASE_DEV_ACCESS_TOKEN, which can delete the project, and SUPABASE_DEV_DB_URL, a direct login that ignores every permission rule over an UNMASKED copy of real student records. Every consumer of those two is migration tooling; a contributor editing src/ needs neither. They were handed to everyone because they arrived in the same block and the guard asserted all four.
Fix. tools/dev-env.mjs maps the dev pair into what the app reads, prints which database it chose, and sets VITE_ENV_NAME so the existing ribbon paints. Both vite configs call it only on command === 'serve' — a production build must never be repointed, and that gate is asserted. .env.local.example now ships the powerful pair commented out; env-check.mjs REQUIRES two and merely reports the other two, so a correctly-provisioned contributor gets a green check instead of being told their setup is broken.
Where it lives now. tools/dev-env.mjs, vite.config.js, passport/vite.config.js, tools/env-check.mjs (REQUIRED / OPTIONAL), and src/js/dev-env.test.js — which asserts the PROPERTY: feed in the file the guide tells a person to create, and a real database must come out. Both bugs were reintroduced and each failed on its own assertion before restoring.
A third fix, from the owner, the same day. Told about the split, they went straight past it: "i think making the contributor cp .env.example, input each key manually is bug prone, i thought of handling the file, or someway automate". Right, and it makes the first two fixes cheap to hold: npm run setup (tools/setup-env.mjs) takes the message a maintainer sent, pasted whole — greeting, code fence, export, quotes, CRLF, and a key a chat client wrapped onto two lines — and writes the file. The transcription step is gone, so the three failures the guide used to list cannot happen. It MERGES rather than overwrites (a maintainer's .env.local holds production credentials and the VM sudo password), backs the file up first, never prints a value, and REFUSES a paste containing a production name.
⚠️ And its own first bug was in the half the unit tests could not see. Nine parsePaste cases were green while the real command was broken: the READ loop ended the paste at the first blank line after any non-empty line, and people send a covering note — so "hey, here you go" + blank closed input before one credential arrived, and it then told the reader they had pasted nothing. Found by piping a realistic message through the actual command, not by reading it. opensAssignment is now the terminating condition and is asserted directly.
A fourth fix, from the owner pressing the same point further. "incase in the future there's more key, or key is changed, it would be tiresome to manually copy paste each key" — the toil is not the typing, it is that nothing told anyone. A variable added to the project reached a contributor's machine only when something broke in a way that did not mention it.
REQUIRED and OPTIONAL were two hand-written arrays beside the file they described — class 6, and the reason a new variable was invisible. They are now DERIVED from .env.local.example (tools/env-manifest.mjs): an active NAME= line is required, a commented # NAME= line is database-work only. One edit to that file and env:check starts asking, setup starts accepting, env:share starts offering, and every contributor's next npm run dev names it and says what to ask for — silent when nothing is wrong. Proved by adding a variable to the example and watching all four react with no code change.
npm run env:share closes the sending side: the maintainer stops hand-picking lines out of a file that also holds SUPABASE_DB_URL and SAMO_VM_SUDO_PASSWORD. It cannot emit a name the example does not offer a contributor, and refuses to run into a pipe without --force.
⚠️ And that refactor broke the whole suite in a way that named no cause. Importing env-check.mjs from vite.config.js put its #!/usr/bin/env node mid-bundle — Vite bundles its config — and esbuild answered tools/env-check.mjs:1:396: ERROR: Syntax error "!" for a file whose line 1 is 19 characters. Fixed by moving the shared part into a shebang-free env-manifest.mjs; dev-env.test.js now walks the config's import graph and fails with a sentence instead.
A fifth fix — the owner asking the question that invalidated my answer."but then i have to be access to my computer to can do it" and "i would have to send it many times for each person isn't it". Both true. env:share automated the maintainer's TYPING and left them as the DELIVERY MECHANISM — once per person, for ever, laptop required. npm run env:pull removes them: values live in one vault item, each contributor fetches their own.
⛔ AND I HAD TOLD THEM THAT WAS IMPOSSIBLE, from a search result rather than a test. I wrote in HANDOFF that the Bitwarden CLI "expects a bare HTTPS root" so our /vault/ subpath ruled it out. One command disproved it:
bw config server https://samo.md.kku.ac.th/vault → Saved setting `config`.
bw login <nonexistent user> → "Username or password is
incorrect" (it REACHED
the identity endpoint; a
wrong URL gives a
CONNECTION error)The vault had been publishing the answer at /vault/api/config the whole time. A search result about a tool is not a measurement of that tool against YOUR deployment, and writing it into the handoff as a constraint would have closed off the right design for however long the note survived.
⚠️ A REAL BUG, found only because a test for the new tool exercised the old one. parsePaste continued a wrapped value onto any line without whitespace — so the ───────────── rules that env:share prints around its block got glued onto the end of the anon key. Pasting exactly what you were shown produced a key wrong by thirteen invisible characters, surfacing later as "the database refused the key (401)". Fixed by requiring a continuation line to be token characters AND to contain at least one alphanumeric — * and - are legal inside a password, so the charset alone could not do it, but a line of PURE punctuation is never half of a credential. Nine earlier parsePaste cases were green throughout: none of them pasted the tool's own output.
The general rule. A guard that compares a list to a list proves the two lists agree, not that either one works. Here both lists were right and the thing between them did not exist. Ask what the lists were meant to PRODUCE and assert that instead — the property, end to end, from the artefact a real person creates. And when a setup hands over N credentials, ask what each one opens: "the SUPABASE_DEV_* block" is a name for a group, and a group name is how a project-deleting token and an RLS-bypassing database login end up on a volunteer's laptop because they wanted to change a colour. The corollary for deciding how to DELIVER secrets: settle what is in the envelope before choosing the envelope — and then delete the step where a human retypes what is inside it. A pure function's tests do not cover the loop that feeds it: exercise the COMMAND against a realistic input, or the half you did not extract stays unproven. And the most realistic input to a paste-parser is the output of your own tool — the one shape nine hand-written cases all missed.
npm run migrate:status --dev answered about PRODUCTION for months
Symptom. Found on 2026-09-06 when the owner said "i want you to write docs properly, check if you've write properly" — an instruction to VERIFY rather than assert, which is the only reason this was looked at.
README.md told people to run:
bash
npm run migrate:status --devCause. npm treats --dev as its OWN flag and never passes it to the script. Measured, same line, one difference:
npm run migrate:status --dev → project fheueuowbchsnsvbcgil [PRODUCTION]
PENDING: 0
npm run migrate:status -- --dev → project xibugtlsphcfuvstnxxh [samo-dev]
PENDING: 3A wrong answer wearing the right question — and not a quiet one: it reports PRODUCTION is in step to somebody who believes they asked about samo-dev, while dev really has three pending migrations. The tool is innocent; it prints [PRODUCTION] precisely so this is catchable, which is exactly the guard docs/state/HANDOFF.md §7 already told readers to rely on. The DOC handed them the broken command.
The same shape had just been introduced by me in three fresh places (npm run env:share --db, --only, --copy), where the failure is quieter still: --db silently yields the two-value block, so a maintainer would believe they had sent four values and sent two.
Fix. Every npm run <script> --flag in the repo's markdown rewritten as npm run <script> -- --flag, with the reason inline at the two places a reader is most likely to copy from. src/js/docs-commands.test.js sweeps all markdown for the broken form, and separately asserts that every npm run <name> on a getting-started page names a script that actually exists in package.json. Both were reintroduced and each failed on its own assertion before restoring.
Where it lives now. README.md, skills/build-the-dev-database.md, skills/onboard-a-contributor.md, docs/start/sharing-credentials.md; guarded by src/js/docs-commands.test.js.
The general rule. A command in a document is code, and nobody ever runs it. Prose is reviewed by reading; a command is only ever verified by execution, and the gap between "this looks right" and "this does what it says" is where a -- lives. Run every command you put in a doc, and diff its output against what the sentence around it claims — here the claim was "the same question, against samo-dev" and the output said [PRODUCTION] in the first line. Same lesson as the deploy pipeline that discarded its failing step: the instrument was printing the answer all along and nobody was made to look at it.
env:pull repointed the user's personal Bitwarden CLI, and its guard was satisfied by a comment
Symptom. None — nothing failed. Found on 2026-09-06 during an audit the owner asked for ("keep checking until you find no error"), by RUNNING a command whose failure path had only ever been reasoned about.
Cause 1 — a project tool wrote outside the project. tools/env-pull.mjs called bw config server <our vault> with no appdata override, so the Bitwarden CLI wrote to ~/Library/Application Support/Bitwarden CLI/data.json and silently repointed a personal bw at the SAMO vault. Confirmed by finding the file the run had just created, with "base": "https://samo.md.kku.ac.th/vault" in it. Anyone who uses bw for their own passwords would have found it talking to us, with nothing anywhere to say why.
Fix 1. BITWARDENCLI_APPDATA_DIR pinned to a gitignored .bw/ inside the project, set in the ONE helper every invocation goes through, so it cannot be forgotten at one call site. The file the audit itself created was deleted, so the machine was left as it was found.
Cause 2 — and this is the more useful half. The guard written for Fix 1 asserted expect(SRC).toMatch(/BITWARDENCLI_APPDATA_DIR/) against the RAW file. The fix ships with a JSDoc paragraph explaining the hazard, and that paragraph contains the name. So deleting the actual override left the test GREEN — verified: reintroduced the bug, ran the test, 14 passed.
That is "satisfied by PROSE" in .claude/rules/mistakes.md class 7 — the same shape as confirm-modal.test.js matching a comment — and the only reason it was caught is the ritual: reintroduce the bug and watch it fail on the assertion you expect. It did not fail, so the guard was wrong.
Fix 2. Read through stripComments(), the shared instrument, plus a control asserting the stripper is actually removing the explanatory comment — so if the stripper ever stops working, the guard says so instead of quietly reading prose again.
Where it lives now. tools/env-pull.mjs (the bw() helper and BW_DIR), .gitignore, src/js/env-pull.test.js.
The general rule. A guard that greps a source file is reading the comments too, and comments describe the fix in the fix's own words — which is exactly the vocabulary the assertion uses. Strip comments before asserting on source, and control that the stripping happened. And the wider one: the failure path you have only reasoned about is not tested. Running npm run env:pull once, in the state every contributor is actually in, took under a minute and found a side effect on somebody else's machine that no amount of reading would have shown.
npm run env:pull printed "Signing in." and then nothing, for ever
Symptom (owner, 2026-09-07, the first real run of the whole path):
"i only see these text, nothing happens"
Fetching your credentials from https://samo.md.kku.ac.th/vault
Signing in. Use your SAMO vault email and master password.
(Nothing is stored in this project — the session ends when
this command does.)and then the cursor sat there. No error, no timeout, no CPU. Every theory it invites is about the network or the vault — a slow TLS handshake, the /vault/ subpath, a 17 MB npx download — and all of them are wrong.
Cause. bw writes its prompts to stderr, and bw() piped stderr. stdin was inherited, so the CLI was genuinely waiting for an answer to a question that had been captured into a buffer nobody read. Measured, which is the only reason this took minutes instead of an afternoon:
$ bw login --raw </dev/null 1>out 2>err
out: (empty) ← --raw puts the SESSION KEY here, which is why
stdout was captured in the first place
err: ? Email address: ← the promptThe capture was not careless: --raw prints the session key on stdout and the tool must read it. stderr was piped alongside it out of symmetry, and because die() quotes stderr to say what went wrong. Both are right for the six subcommands that never ask anything, and fatal for the two that do.
Second, smaller cause found in the same run: bw login exits 0 with empty stdout when its prompt reaches end-of-input. So the run continued with BW_SESSION='' and died four steps later at get item, blaming the vault for a sign-in that had never happened — the failure naming the wrong step.
Fix. stdioFor(args), exported so it can be asserted: stderr is 'inherit' for login/unlock, 'pipe' for everything else, stdout always captured. Plus an explicit if (!session) that fails at the step that actually failed, and a line before the first call saying the first run downloads ~17 MB — that call's progress is piped away too, and it is the other silent wait.
The guard, and what it caught on the way. src/js/env-pull.test.js asserts the property both ways: prompting commands must inherit stderr, non-prompting commands must capture it (that half is not decoration — inheriting everywhere would throw away the diagnostic die() quotes). Reintroduced the bug, watched it fail on that assertion, restored.
Writing it turned an OLDER guard red: "locks the vault on every exit path" counted occurrences of the STRING --raw in the raw file, so three sentences of a new comment explaining what --raw does broke it while the code was correct. Same file, same lesson as the entry above, one entry later — and the fastest way to green was to edit a number. It now counts CALL SITES in the stripped source, with a control that fails if it has no subject.
Where it lives now. tools/env-pull.mjs (stdioFor, the !session guard), src/js/env-pull.test.js.
The general rule. A tool that asks a question on a stream it has captured is a hang, and it looks exactly like a network stall — which is what everyone debugs first, away from the code. Before capturing a child process's output, ask whether that child ever needs to ask the human something; the streams that carry QUESTIONS and the streams that carry DIAGNOSTICS are not the same set, and stdio is one decision for both. Class 7's "the instrument can delete the witness", moved one step earlier: here the deleted witness was not the evidence of a failure but the prompt that would have prevented it.
CI was red for 19 consecutive runs, and nobody read the last one
Symptom. A push comes back Status: Failure — 2 failures · 1842 passes. The two are in dev-env.test.js, and locally the same suite says 1848 passed. The obvious reading is that the commit just pushed broke something.
It did not. build.yml had failed on every push since 2026-09-06 04:53 — 9aab626 through e90bbc8 is eighteen runs across 33 hours, and the push that prompted this was the nineteenth, indistinguishable from the eighteen before it.
Cause. Three assertions passed the maintainer's OWN .env.local:
js
applyDevDatabaseEnv(procEnv, join(ROOT, '.env.local'));That file is gitignored because it holds production secrets, so it exists on every maintainer's laptop and can never exist on a CI runner. The guards were green where nothing was at stake and red where the check is actually enforced.
Two things made it worse than an ordinary red build:
- One of the three was green on CI for the reason it exists to catch.
expect(env.VITE_ENV_NAME).not.toBe('production')passes onundefined, andundefinedis precisely "no ribbon, looks like production". So the CI signal was two failures where the honest count was three. - A build that is red for a reason nobody can fix stops being read. Nothing was ignored on purpose; local
npm test— the repo's standing check before every commit — kept answering1848 passed, which is the answer to a different question than the one CI asks.
Fix. contributorEnvFile() writes the existing in-memory contributorEnv() fixture — built FROM .env.local.example, so it tracks the contract — into a temp dir, and the three assertions take that path. The guards now assert the same property with no secret present, which means they assert it everywhere.
The ratchet. A sweep over every *.test.js in the repo asserting none of them reads this repo's real .env.local, with a control that fails if the walk finds no files. It runs in the normal suite, so the next instance is caught on the laptop rather than 19 pushes later. Reintroduced one of the three original call sites, watched it go red, restored.
Where it lives now. src/js/dev-env.test.js.
The general rule. A guard that needs a secret cannot run where guards are enforced. Anything a test reads must be either committed or synthesised — feed it the artefact a real person creates, never the one your own machine happens to have. And the meta-rule this cost 19 runs to learn: npm test passing locally is not the same claim as CI passing, so when a CI failure names tests that pass locally, the first question is not "what did I break" but "how long has this been red" — gh run list --workflow=build.yml answers it in one command.
The first migration replay condemned 27 healthy migrations, then found the truth
Symptom. A new CI job replays all 180 migrations onto an empty Postgres. Its first run stopped at 0153 with:
ERROR: relation "_flatten" does not exist0153 has been live in production for weeks. Read as written, the job was saying the database cannot be rebuilt from this repo — and it would have said the same about the 27 migrations after it, none of which it ever reached.
Cause 1 — the instrument, not the subject. 0153 does the honest thing for a data conversion:
sql
create temp table _flatten on commit drop as ...
update public.team_nodes n ... from _flatten f ...on commit drop means the temp table dies when the transaction ends. tools/apply-migration.mjs POSTs a whole file as one query, so the file is one transaction and _flatten survives to the last statement. psql in its default autocommit makes every statement its own transaction — so the table was dropped the instant it was created. The replay was faithfully testing psql's transaction semantics and calling the answer a migration bug. --single-transaction makes it test the path that actually runs in production.
Cause 2 — a refusal is not a break. With that fixed it reached 0166, which backfills timelines and then checks its own work: if events < 300 then raise exception '0166: only % timeline events left'. An empty database has 0, so it refused. Correctly — it is a data migration and there is no data, which says nothing about the schema.
The tempting fix is a list of file names to skip. That is the exemption that outlives its reason (class 7): it rots on the first rename, and it covers every future failure of that file, including real ones. The discriminator used instead is structural — the SQLSTATE. P0001 is raise exception, a human refusing on purpose; a missing column is 42703, a missing table 42P01, a syntax error 42601. P0001 is reported as refused and the run continues; everything else stops it. It needs VERBOSITY=verbose, or psql prints no SQLSTATE and every error looks alike.
The guard that matters is the control. classify() is asserted in both directions: P0001 refuses, and 42703 / 42P01 / 42601 / a message with no SQLSTATE at all are NOT waved through. Without that second half, "refused" could quietly widen to mean "any error" and the job would go green on a broken schema — which is worse than not having it.
The result, and how it was checked. 180 migrations apply to an empty database in ~9 s, producing 65 tables, with 0166 refused. The real samo-dev reports 66 (53 public + 13 passport) — and the extra one is _timeline_backup_0166, the table the refused migration creates. The counts close exactly, which is the difference between "the job exited 0" and "the schema it built is the schema we have".
Where it lives now. .github/workflows/migrations.yml, tools/replay-migrations.mjs, tools/ci/supabase-platform.sql, src/js/replay-migrations.test.js.
The general rule. Before believing a proof's verdict about your code, check that the harness runs your code the way production does. Transaction boundaries are the classic gap — the same SQL is correct in one and broken in the other, and neither the file nor the error mentions the difference. And when a proof must forgive something, forgive a property the database reports (a SQLSTATE, a class of error), never a name you typed in: a name list is a second copy of a decision, and it goes stale the way every second copy does.
A throwaway branch deleted two files' worth of work, and the commit message covered it up
Symptom. docs.yml had not run for a commit whose message described changes to two files under docs/. The workflow triggers on docs/**, so either the trigger was broken or the commit did not contain what it said.
Cause. The commit contained one file. The two documentation edits had been made, verified (docs:build, full suite) and left uncommitted while a throwaway branch was created to test something unrelated:
git checkout -b ci/probe-… # working tree was dirty
git add -A && git commit # swept the doc edits in with the probe
…
git branch -D ci/probe-… # and deleted themgit add -A does not distinguish the change you are testing from the change you happened to be holding. The later commit's message was written from the session's memory of the work rather than from its diff, so it confidently described files that were no longer in the tree.
Fix. The dangling commit was still in the object store, so git checkout <sha> -- <paths> recovered the exact reviewed text rather than retyping it, and the probe's own file did not come back with it. The message of the commit that lied is corrected in the message of the one that restores it — history is not rewritten here.
Where it lives now. Nothing to guard in code; the discipline is below.
The general rule. A throwaway branch created from a dirty working tree takes the dirt with it, and deleting the branch deletes it. Commit or stash BEFORE branching for an experiment. And the detection lesson, which is the more transferable half: a commit message that names files is a claim, and --stat is the check — this one was caught only because a path-filtered workflow did not fire, which is a strange thing to notice and easy to miss. When a message says "also changed X", read the diffstat before believing it.
A registry sweep that reported CLEAN over a policy planted in the same transaction
Symptom: proj0181-prof-upload.sql grew a section E — "which SELECT/ALL policies can only answer by consulting the table they protect?" — and it printed the one expected entry, PASS. It looked authoritative for an hour. Then the break-it ritual planted a second, deliberately self-referential policy inside the same rolled-back transaction, and E1 still said PASS.
Cause: pg_get_expr() renders a policy's expression with SQL keywords in UPPERCASE and the schema qualifier stripped —
(EXISTS ( SELECT 1
FROM shop_orders o
WHERE ((o.id = shop_order_items.order_id) AND …)))— and the sweep matched expr ~ ('from\s+(public\.)?'||tbl). ~ is case-SENSITIVE, so the inline half of the sweep had never matched anything, ever. The one entry it did report came from the OTHER half (a function whose body reads the table), which happened to work. Half a working instrument reads exactly like a whole one when the answer it gives is the answer you expect.
Fix: ~* for both halves, and the ritual repeated until the planted policy made E1 go red. The corrected sweep still returns the same single entry on production — but that conclusion is now earned rather than assumed.
Where it lives now: tools/proj0181-prof-upload.sql §E, with the reason written at the regex.
The general rule. When you grep a DATABASE's own rendering of something, grep what the database PRINTS, not what you typed. pg_get_expr, pg_get_functiondef and pg_get_triggerdef all re-render from the parse tree: keywords upper-cased, schema qualifiers dropped, whitespace normalised, != rewritten to <>. A pattern written from the migration source will silently miss. And the second half, which is the one that nearly got away: a sweep whose result MATCHES YOUR EXPECTATION is the hardest kind to doubt — plant the thing it is supposed to find and watch it go red, because "it returned what I thought" is not evidence that it looked.
Repairing our own damage: the fix that bills the user is the wrong fix
Symptom: three signed PDFs reached Google Drive while the database row was refused (0181), so the app showed อนุมัติแล้ว with no signature. The first remedy proposed was "ask the อาจารย์ to upload them again". The owner rejected it: "it's web fault… it isn't client fault, do it the best practice way."
Why it was wrong. The artefact was never lost — only our record of it. Asking the person who already did the work correctly to redo it bills our defect to them, creates a SECOND Drive file while orphaning the first, and stamps the history with today's date, erasing the fact that they signed on 3 September.
Fix: tools/proj0181-repair-orphans.mjs re-attaches the EXISTING file, dated the approval instant, attributed to the อาจารย์ the request named, with a timeline entry on both request and document saying plainly that a system repair happened and why. Dry-run by default; idempotent on drive_file_id.
Two things it does that a quicker script would not:
It refuses to invent a timestamp. Drive's own creation time is exposed by no handler this project has, so it uses the APPROVAL time and the note says that is what it is. A plausible invented time would have sat in the audit trail forever, unquestioned.
It does not trust the ids it was given. Each file is read back through GAS and must pass five checks — is a PDF, not already attached, not the original itself, LARGER than the original, same filename. All three were bigger than their originals with a different PDF producer version (1.6 vs 1.3/1.5), which is what signing and re-exporting produces. The filenames independently confirmed the mapping, so it never rested on the order they were pasted in.
The general rule. When a bug of ours destroys a record, repairing the record is our job, not the user's — and a data repair is held to a HIGHER evidentiary bar than a feature, because it writes history that nobody will re-derive. Verify every input against the system of record, never invent a value to fill a column, and say in the data itself that a repair happened.
A control reported the detector blind, and the detector was fine — an earlier step of the same proof had deleted its subject
Symptom. authz0182-insert-returning-seam.sql, first run: 20. detector FLAGS the deliberately bad policy → (nothing — DETECTOR IS BLIND). Ten other assertions passed, including the whole live reproduction of the bug above it. The obvious reading — the pg_depend detector cannot see a self-lookup policy, so the sweep in §30 is worthless — was wrong.
Cause. The proof has two jobs and used one table for both. §A demonstrates the mechanism: create a table with a self-looking-up SELECT policy, show that a bare INSERT is allowed while the same INSERT … RETURNING * is refused, then drop that policy and install the fixed one to show RETURNING start working. §B then ran the detector and asked it to find… the policy §A had just dropped. The detector answered correctly about a database that no longer contained the thing it was asked about.
Fix. A second table, _seam_detect, that carries the bad shape and is never rewritten, so the mechanism demonstration and the detector's control do not share a subject. Both are created and rolled back inside the same transaction.
Where it lives now. tools/authz0182-insert-returning-seam.sql §A2, whose comment says why the table exists, because the reason is not guessable from the code. The detector now flags it (§20) and still does not flag the healthy two-argument prof_can_see_file (§21) — the two directions that separate "no policy is broken" from "this cannot see a broken one".
Two general rules, and the first one is the reason to keep writing controls.
A control that fires is doing its job even when the thing it indicts is innocent. This one read as an indictment of the detector and was really an indictment of the proof's own setup. The instinct on seeing DETECTOR IS BLIND is to go and fix the detector; the correct first question is whether the scenario still contains what the assertion is looking for. A proof with no control would have printed ten greens and shipped a sweep that had never been shown to see anything.
A proof that MUTATES its own fixtures needs one fixture per claim. The moment a file both demonstrates a broken state and demonstrates the repair, any later assertion about the broken state is reading a world that was edited out from under it. This is the same shape as the entry above where a proof deleted only its own samples and went green while the reporter was paused: the setup and the assertion disagreed about which world they were in, and only the ORDER of the statements said which one won. Give each claim its own subject, or assert it before you repair anything.
A --rows-only flag turned the sweep's own anti-vacuity rule against it: it printed "0 orphans" for a direction it never examined, then said "in both directions"
Symptom. node tools/proj0183-drive-orphans.mjs --rows-only finished with 123/123 rows stat'd · 0/63 folders listed and then printed:
✓ Drive files with NO database row (the 0181 shape): 0
…
✓ Drive and the database agree, in both directions.Every one of those zeros was true of the half that ran and meaningless for the half that did not. The last line was simply false.
Cause. The tool was written with exactly this hazard in mind — its header says "'No orphans found' and 'the listing failed' must never render the same" — and it guards it with a control that FAILS when listed === 0. Then --rows-only was added so the cheap direction could run without pressuring a shared endpoint, and that control had to be suppressed for the deliberate case:
js
if (!ROWS_ONLY && listed === 0) { /* fail */ }The suppression silenced the CONTROL and left the CLAIM. Nothing else in the reporting knew a half had been skipped, so an un-run direction rendered identically to a clean one — which is the precise failure the control existed to prevent, now reachable by a supported flag.
Fix. show() takes a ran argument and prints — NOT EXAMINED (this run skipped that half) instead of a count; the examined line says folders: SKIPPED rather than 0/63; and the verdict names the directions that actually ran, with an explicit ⚠ THE OTHER DIRECTION WAS NOT EXAMINED. This is not a clean bill of health for it. Re-run and read before trusting.
Where it lives now. tools/proj0183-drive-orphans.mjs, in show() and the verdict, with the comment explaining why the argument exists.
The general rule, which is new and worth the entry. When you add a flag that legitimately skips part of a check, the control you must suppress is the least of your problems — every SUMMARY LINE downstream is now able to lie. A guard's output is a claim about a scope, so narrowing the scope invalidates the claim, not just the assertion. The tell is a suppression written as if (!FLAG && <original condition>) — that edit acknowledges the flag in exactly one place, which is proof the rest of the code still believes the full run happened. Grep for everything that reads the same variable the control read — here listed — and make each of them say "skipped" rather than "zero". And a tool that guards against a vacuous pass is not exempt from producing one; this one had the rule written in its own header, in capitals, and still did it.
A brand-new guard was green on the laptop and red on CI within an hour — it read .env.local at import time
Symptom. src/js/projects/drive-orphans-report.test.js passed 13/13 locally, was committed and pushed, and CI's build job went red with all 13 failing — both does not crash, --rows-only stats its rows, every one. The tool it spawns had not changed between the two runs.
Cause. tools/proj0183-drive-orphans.mjs read the maintainer's gitignored .env.local unconditionally at module top, before anything could branch on SELFTEST. The self-test needs no credential — sql() and gas() are both stubbed — but the readFileSync threw ENOENT on CI, where that file does not and must not exist, so the process died before the reporting block it exists to test could run.
Fix. Read the file through a readEnvFile() that returns '' on ENOENT. The credential CHECK stays exactly where it was, so a REAL run still refuses with need VITE_SUPABASE_URL + SUPABASE_ACCESS_TOKEN in .env.local; only the self-test path stops needing it. Verified by reproducing CI's condition rather than trusting the reasoning:
bash
mv .env.local /tmp/ && npx vitest run src/js/projects/drive-orphans-report.test.js # 13 passed
node tools/proj0183-drive-orphans.mjs --rows-only # refuses, names the vars
mv /tmp/.env.local .Where it lives now. The tolerant read carries the reason in a comment, and the test file opens with the mv .env.local recipe so the next person verifies it the way CI will.
The general rule — and this repo had already paid for it once, in the same file. A guard that needs a secret cannot run where guards are enforced. The earlier instance kept CI red for 19 consecutive pushes while local npm test said 1848 passed. This one was caught in an hour only because CI was checked; nothing else would have said so, since the local suite is green by construction on the machine that owns the credential.
What makes it recur is that the dependency is INVISIBLE at the assertion. None of these 13 assertions mentions a credential; the dependency lives in a module-level readFileSync inside the thing under test, three files away. So the question to ask is not "does my test need a secret" but "does anything my test LOADS read one at import time" — and the only reliable way to answer it is to move the secret aside and run. Do that before pushing a test that spawns or imports a tools/ script. And when a brand-new guard goes red on CI, suspect the guard's environment before the code it guards.
Two live proofs went red, and STATE.md recorded the wrong cause — the scenario borrowed its geometry instead of building it
Symptom: npm run proofs reported 33 of 35 green for days. claude0154-quota-guard and claude0155-free-now were the two red ones, and STATE.md explained them as "Both must BOOK a slot; the live 5-hour window is already claimed so claude_booking_guard refuses" — the "scenario needs live geometry that RAN OUT" trap — with the instruction "make the proof CREATE the geometry".
Cause: that explanation was a HYPOTHESIS and it was wrong on the mechanism. Running 0154 printed 19 of 20 passing — not a proof that could not book, but a single failing case, A5. a booking that exactly fills the week is allowed. Nothing was refusing a slot; the WEEKLY POOL was short. On 2026-09-07 the first real booking anybody has ever made (แก้ไขระบบActivity score ภายในสาขา, 70%) landed inside the very week both proofs had hardcoded as their quiet stretch. 0154's arithmetic comment says "30+70+1+50 = 151 booked so far" and then squeezes week_pool_pct to 161 — true only while the week holds nothing else. Both files stated the assumption in a COMMENT: "clear of anything real".
The expensive half is what the same bug was doing in the OTHER direction. 0155's C3. a still-to-run block stays reserved, whenever you ask was PASSING, and it was passing on the real user's booking: it expected reserved_pct 70 and the probe's own block had already been released at that instant, so the 70 it read was the stranger's. Clearing the week turned C3 red — a false green exposed by fixing an unrelated case. 0155 had three stale assumptions, all in comments: "clear of anything real", "sampled_at … months after every real row" (real sampling every 15 min had long overtaken the hardcoded date), and "the scenario here is entirely in the future" (it was six days past).
Fix: each proof now CONSTRUCTS the absence it needs and then ASSERTS it, inside the begin … rollback both files already end with — the same pattern 0155 was using on claude_usage_samples all along. A delete scoped to the probe week, then a control case A0 that fails with the actual number on the page. 0155's anchor is no longer a date somebody typed: it is claude_week_start(now() + interval '28 days') + interval '2 days 8 hours', which is a property (a Saturday in a future quota week, later than every real sample and later than now) rather than a date that was briefly that property. Both are green: 0154 21/21, 0155 23/23, suite 35/35. Each control was falsified before being trusted — remove the delete and A0 reports got=70 beside A5; inject a stranger's booking into the future week and A0 reports got=25 at the top of the verdict.
Where it lives now. tools/claude0154-quota-guard.sql and tools/claude0155-free-now.sql, each with the reasoning above the delete.
The general rule. A scenario that needs an ABSENCE must create it and assert it; a comment claiming the environment is empty is the tell. This repo already had that rule from claude0167 and it did not stop the next instance, because the assumption reads as scene-setting rather than as a dependency. Two sharper tests for the next reader:
- A date somebody typed is not a property. "Clear of anything real", "months after every real row" and "entirely in the future" were all TRUE when written and all decayed on their own. If a scenario needs a property, derive it from
now(). - Suspect a passing case that sits beside a failing one in the same arithmetic. C3 was green because a stranger's row happened to carry the same number the proof expected. When you fix the environment, re-read what went red — a case that newly fails was never really passing.
And do not trust a recorded cause over a run. STATE.md's diagnosis had no — how, cost nothing to check, and sent the reader at the booking guard instead of the weekly pool. One node tools/db-query.mjs printed the real answer in seconds.
A redaction written against the healthy shape prints the secret on the broken one
Symptom. The owner asked "check if i've done it correctly i think i don't" about /etc/samo-discord-bot.env on the VM. The inspection script was written to print structure and never the value — mode, byte count, value length, a masked first-and-last-four. It printed the entire bot token into the chat transcript. Third leak of a Discord token into a transcript for this project (2026-08-28, 2026-09-11, 2026-09-12), and the first one caused by the tool built to prevent it.
Cause. Every redaction in that script keyed on the shape it was testing for. The last line was:
bash
sed -n '2,$p' "$f" | sed -E 's/=.*/=<redacted>/'which redacts everything after an =. The file's actual defect was that the token had been pasted bare, with no DISCORD_TOKEN= prefix — so there was no = on the line, the substitution matched nothing, and sed passed the line through unchanged. Every other field in the same script reported 0 or empty for the same reason and looked like a tidy, safe result; the one field with a fallback-to-raw was the one that fired.
The masking was conditional on the file being well-formed. The whole point of running it was that the file might not be.
Fix. The instrument must be incapable of emitting content, not merely instructed not to. The replacement reads the file line by line and prints only length, blank, and a yes/no for whether the line starts with the expected key — there is no code path in it that echoes a line, so no input can make it leak. It diagnosed the same file correctly: a 72-character bare value on line 2, no key anywhere, four blank lines.
Where it lives now. The token was reset. .claude/rules/security.md's DISCORD_TOKEN row carries the third leak date.
The general rule. A redaction conditional on well-formed input is not a redaction — it is a guess that fails open exactly when you are looking at the malformed case. Never write mask(x) as "strip the part after the delimiter"; write the reporter so that raw content has nowhere to go — emit derived facts (lengths, counts, booleans, hashes) and never the string. Test it against the broken input, not the healthy one: a masker verified only on a correct file has been verified on the one input that was never going to leak. This is class 7's instrument trap in its sharpest form — the guard fails green on the healthy case and fails open on the case it exists for — and class 2 underneath it, an unresolvable reference answering "allowed".
team0183 had the "describes today's data" defect TWICE, and one of them was 42 ticks from a false red
Symptom. None yet — both were provoked deliberately. skills/discord-role-sync.md carried a standing note that team0184 had gone 18/18 → 16/2 because a provisioning run succeeded, and that "the other two plausibly have the same shape somewhere and nobody has provoked it". Provoked on 2026-09-13.
Defect 1 — an assertion that measured a RATIO of live data the owner is being asked to change.
sql
insert into probe select '51. …and left the majority untouched', 'true',
(select (count(*) filter (where not discord_role) > count(*) filter (where discord_role))::text
from public.team_nodes);Measured: 107 of 299 ticked, so it flips at 150 — 42 more ticks. And STATE.md says in as many words that the remaining nodes are the owner's review. So the owner doing exactly the work asked of them turns this red while nothing is wrong, and the fastest route back to green is to edit the comparison, which is how a guard stops meaning anything.
The comment above it stated the real rule — "if a later edit turned the seed into 'tick everything', 51 would go red" — and the SQL did not implement it. It now asserts that rule directly (some node remains unticked), with a control over a synthetic all-ticked tree, because provoking the real thing would mean updating 299 live rows through a permission-recompute trigger in order to roll them back.
Defect 2 — a scenario FOUND in production rather than created, which production is being asked to remove. §22-24 need two team_nodes sharing a name, and took them from whatever duplicate production happened to hold. HANDOFF §14b item 2 asks the owner to rename or untick five contested ฝ่าย — i.e. the proof's subject is on somebody's cleanup list. It now prefers a real pair (the shape the application actually produces) and BUILDS one when there is none.
⚠️ And the claim written into the fix was wrong until the run corrected it. The comment first said a duplicate-free world would make §22-24 pass VACUOUSLY — an UPDATE matching no rows raises nothing, so "22 answers ok having changed nothing". Forcing that world showed the opposite: 4 FAIL, with 22/23/24 all answering deny-rls, because this proof's own pg_temp.attempt() scores a zero-row UPDATE as deny-rls rather than ok — the "three answers, not two" instrument in its preamble doing precisely its job.
That is still a defect, in its other costume: the proof fails saying deny-rls, sending the next reader after a row-security problem that does not exist when the truth is that the scenario ran out. A misdiagnosis costs more than a plain failure. But the write-up had to be corrected to what was measured, not what was reasoned.
A third thing the fix exposed: the expected count had FOUR homes. 22 was written in STATE.md, docs/DISCORD-ROLE-SYNC.md, docs/state/HANDOFF.md and skills/discord-role-sync.md. Adding one assertion meant correcting all four — which is how one of them stays stale and starts lying. The skill, the operational home where a wrong number does the most damage, now carries no count at all: every row must say PASS.
Where it lives now. tools/team0183-discord-mapping.sql §B and §E; skills/discord-role-sync.md.
The general rule. When a proof's subject is on somebody's to-do list, the proof is scheduled to break. Both defects here read live state that a HUMAN has been explicitly asked to change — the tick-box the owner is reviewing, the duplicate names the owner is renaming — so "green today" was a statement about how far through their task they were. Before trusting an assertion over production data, ask who is allowed to change this, and has someone been asked to? If the answer is yes, assert the rule instead, and construct whatever geometry the scenario needs. And never let a proof's expected count live in more than one place; better, let it live in none, and require every row to pass.
Two more proofs whose SCENARIO had run out — and one that asserted "the first trigger, whichever it is"
Symptom. npm run proofs on 2026-09-13: 2 of 39 not green, neither caused by the change being made.
shop0150-buyer-contact.sql ✗ errored: null value in column "id" of relation "shop_orders"
team0185-link-codes.sql ✗ 75. updated_at is maintained by a TRIGGER, not by callersshop0150 — borrowed geometry that ran out, and it ERRORED rather than failed. The proof builds its subject by copying an existing order as a template:
sql
to_jsonb((select o from public.shop_orders o order by o.id limit 1))shop_orders now holds 0 rows (measured). The subselect is NULL, NULL || jsonb_build_object(...) is NULL, and jsonb_populate_record returns a row of all nulls — so the insert violated NOT NULL and the transaction aborted before a single assertion was emitted. Zero rows of verdict, an HTTP 400, and only the runner's r.status !== 0 → FAIL branch stopped that being silence.
Same class as team0183 §22-24 fixed the same day: if the thing a proof needs can run out, create it. The floor is now supplied on the left of the || so a real template still wins when one exists.
⚠️ And jsonb_populate_record does not apply column defaults. A key absent from the jsonb becomes NULL; it does not become def=0. Supplying only the three columns with no default got as far as null value in column "fee". Every NOT NULL column has to be named, defaults included — which is exactly the knowledge the template trick existed to avoid needing, and the reason this failure mode is easy to write.
team0185 §75 — an assertion that described a list one element long.
sql
select p.proname from pg_trigger t … where t.tgrelid = 'public.discord_links'::regclass
and not t.tgisinternal limit 1 -- expected 'touch_updated_at'A limit 1 with no ORDER BY, over a list that happened to hold one trigger when it was written. 0187 added a second one and the proof went red while nothing it guards had changed. Its own comment three lines above says "assert the MECHANISM" — and the mechanism is that a trigger calling touch_updated_at exists, not that it is the only one or that it sorts first. Now an exists (… and p.proname = 'touch_updated_at'), which holds however many triggers the table gains.
The general rule. A proof that reads production state is dated the moment it is written, and the two ways it expires are different. One is the scenario running out — the template row, the duplicate name, the quota week — and the cure is to construct what it needs. The other is an assertion that describes the SHAPE of what it saw — one trigger, two roles, a majority — where the cure is to assert the property that shape was evidence for. Both are found the same way: after any schema change, run every proof, not the ones about the thing you touched. Both of these were in files nobody had edited.
I "corrected" a number that was already right, because I checked its value and not its UNIT
Symptom. docs/state/HANDOFF.md §14b item 7 said "the remaining 51 roles". Re-measuring team_nodes gave 107 ticked, 48 mapped, 59 unmapped, so I changed 51 → 59, wrote a parenthetical saying the old figure predated the adoption run, put it in a commit message, and repeated 59 to the owner in a role-cap table.
It was wrong. 51 was correct. Running the provisioning plan against the live guild the next day:
ADOPT 0 · CREATE 51 · NEAR 4 · CONTESTED 4 51 + 4 + 4 = 59The 59 unmapped ticked nodes split: 51 can have a role created, 4 are near-matches awaiting a human, 4 are contested names. Item 7 is about roles to CREATE — "creating spends 51 of 70 remaining under the cap" — so 51 was exactly the right quantity, and 59 is a true number answering a different question.
Cause. I verified the value and never verified the unit. "Roles still to create" and "ticked nodes without a role" are close enough in English to read as synonyms, and the surrounding sentence — the one that says what is being counted — was the part I did not re-read. The measurement was real, the arithmetic was right, and the correction was still a regression, which is the dangerous shape: it arrives with evidence attached.
Worse, it laundered itself. Once written it was quoted in a commit message and in a table to the owner, so a wrong number acquired three homes in an hour — the same multiplication this repo has already paid for, running in the corrective direction for once.
Fix. Reverted to 51. The lists it sits beside are now re-measured and dated, and the section says explicitly that the plan reports 4 contested where the prose says 5, because a node that already holds a role is counted as already mapped — both right, counting different things.
Where it lives now. docs/state/HANDOFF.md §14b items 2, 3 and 7.
The general rule. Before correcting a number, read the sentence that says what it counts. A measurement can be accurate and still be the wrong quantity, and a correction carries more authority than the thing it replaces — nobody re-checks a figure that has just been "verified". When two counts of the same data differ, the answer is usually that they are counting different things and BOTH belong, labelled; reach for that before reaching for a fix. The tell here was available and ignored: 59 − 51 = 8, and the plan had printed two buckets of 4 directly underneath.
I recommended a cheaper option without measuring it, and it saved one role out of fifty-one
Symptom. Having told the owner that 51 Discord roles needed creating, I added a recommendation: "I'd create only the ones with people beneath them first — that fixes most of the 219 without committing the whole budget."
Measured the next message:
59 unprovisioned ตำแหน่ง
58 have people beneath them
1 emptyThe suggestion saved one role. There was no middle path in that direction, and the sentence had already been read.
Cause. The recommendation was derived from the shape of the problem rather than from the data. "Some ticked ฝ่าย are empty org-chart tidying" was true of the contested list — four of five hold nobody — and I carried that intuition across to a different population without re-asking. It reads as a measured judgement because it sits beside real numbers.
And the real middle path was 50× better, in a direction I had not looked. Most people who are "short a role" still RECEIVE one, from a ticked ancestor that is already provisioned; short means missing the specific role, not missing everything. Only 27 people had no provisioned ancestor at all — and because ancestry is a tree, covering them means creating the node highest in their ancestry, which covers everyone beneath it at once:
27 people ฝ่าย รพ. ร่วมผลิต (one top-level ฝ่าย)One role against fifty-one, and 184 of Discord's 250 cap instead of 234. That is not a refinement of the advice; it is a different decision, and the owner nearly made the expensive one on my say-so.
Fix. npm run discord:readiness now computes and prints the minimum cover itself, so it stays true as ทีม SAMO is edited and nobody has to reason about it again. discord-provision.mjs gained --only '<name>', because the cheap option could be described but not executed — the tool was all-51-or-none, so the recommendation was unimplementable at the moment it was made. skills/discord-role-sync.md now names the bad suggestion explicitly so it is not re-derived.
⚠️ Two bugs were introduced writing that flag and caught before shipping: the narrowing ran AFTER the cap arithmetic that consumed it (a temporal dead zone — the tool would have thrown), and the write loops still iterated the unnarrowed list, so --only would have shown a plan of one role and then created all fifty-one. That second one is exactly the "what gets applied must be what a human read" failure the tool's own count check exists to prevent, arriving through a new door.
The general rule. A recommendation is a claim, and it inherits none of the credibility of the measurements it is printed beside. Before offering a cheaper or smaller option, run the query that says how much cheaper — a fraction, not an adjective. If the tooling cannot execute the option, that is a second reason to check it: an unimplementable recommendation has never been tested against anything. And when the obvious axis turns out to be worthless, look along a different one before accepting the expensive answer — here the useful question was not "which ฝ่าย are empty" but "who would receive literally nothing", and those have wildly different answers on the same data.
A proof whose subjects were "whoever comes back first", three times in one file
Symptom. house0188-unresolved-seat.sql was 24/24 green. Adding one section to it turned two unrelated assertions red and then made the whole script error out with ไม่มีสิทธิ์นำเข้าข้อมูลนักศึกษา — against code that was correct.
Cause, the same one three times. Every subject in the proof was selected as order by id limit N over a live table, so the proof's SCENARIO was whatever the database happened to hold:
- §12 asserted the claimed รหัสนักศึกษา landed on the student row. True only for an actor the registry does not already know — and 0189 made the registry win on that column. The first actor turned out to be a real registered person, so a correct fix made the assertion red.
- §50 asserted a claimed seat is dropped from the held list. It was passing for the wrong reason and hid a real bug (0190) until the actor changed.
- §G called
record_unresolved_rows, which gates oncurrent_user_role()— read from the jwt claims, i.e. whoever spoke last in the script. It had been running as the §F outsider and passing only because the first non-kkumail account in the table happens to be an admin. A new section moved a non-admin into that slot and the permission check fired.
Fix. The proof now CREATES the state each case needs instead of finding it: pg_temp.set_registry() puts a known name and รหัส on each actor's people row (one actor agreeing with the held row, one deliberately diverging), and §G selects an importer subject asserted to actually hold the house grant before speaking as them. Three subjects, each chosen for the property the case is about.
Where it lives now. tools/house0188-unresolved-seat.sql — the §B registry block, importer in §G, §H.
The general rule. A proof that SELECTS its subject from live data is testing the database's current contents as much as the code. "It went green" then means "the first row happened to have the shape I assumed", and the failure arrives later, attached to an unrelated change, pointing at the wrong thing. This repo had already written down the neighbouring version — a scenario needing live geometry that RAN OUT — and the rule generalises: if the case needs a particular state, create it; never take it as found. A subject picked by limit 1 is a subject that changes under you, and the tell is an assertion about a subject's attributes that the proof never set.
The tool wrote both halves of the answer and neither was the file to upload
Symptom. None yet — found while preparing the first real ระบบบ้าน import, one step before doing it. What it WOULD have looked like: the import runs, says นำเข้าเรียบร้อยแล้ว, 1,611 students appear and are correct — and the 165 people the faculty file names but cannot address are simply not there. No error, no warning, right counts everywhere. The damage only surfaces weeks later, as students who cannot find themselves and cannot claim a seat either.
Cause. tools/clean-house-csv.mjs splits a handover file into <base>.clean.csv (the rows that can become students) and <base>.pending.csv (the rows that cannot, each with its reason). Both files are correct. Both are described accurately in the report. The file you would upload — the one holding every line so the importer can act on all of them — did not exist, and the half labelled นำเข้าได้ is the one a reader reaches for.
Uploading it does two invisible things, and both are the system working as designed:
record_unresolved_rows(0188) replaces the held list with whatever the uploaded file could not address — deliberately, so a file that resolves everybody can say so by clearing it. A file that HOLDS nobody is indistinguishable from one that resolved everybody, so all 165 held seats are deleted, andclaim_my_student_seatreads that table.diffAgainstExistingcounts a SKIPPED line as the file MENTIONING that person — also deliberate, so a student who claimed a seat with an address the faculty file has never had is not flagged as gone. Drop the skipped lines and exactly those people get stampedmissing_sinceon the next import.
Neither is reachable from the cleaner's source, which is why no source assertion would have found it: every line of the tool was right, every count in its report was right, and the defect was the ARTEFACT BETWEEN the two files.
Fix. <base>.import.csv — every line of the handover in file order, with the address blanked on exactly the rows the cleaner routed out, so io.js reaches the same verdict the cleaner did and RECORDS it. §0 of the report names it first and says what uploading the other one would do. Line numbers survive: a held row's source_line in the database is the บรรทัด printed in the report.
Blanking a duplicated address has a price, and it is asserted rather than discovered: both holders land as no_kkumail, because by the time io.js reads the file the address is gone. That is the cost of not letting LINE ORDER decide who owns a login — io.js keeps the first and skips the rest — and the report still names both people and the address they shared.
Where it lives now. tools/clean-house-csv.mjs (the .import.csv writer and heldLines), src/js/house/clean-csv.test.js, skills/import-the-house-roster.md. The guard runs the REAL cleaner over a synthetic handover file and feeds the REAL importer; 5 of its 7 assertions were watched going red with the upload file reverted to the clean half. It carries the control that the clean half alone holds NOBODY — the mistake itself, asserted, so the two files can never quietly become interchangeable.
The general rule. A tool that emits several files has also chosen one of them for the reader, whether or not it says so. The repo already knew the neighbouring shape — two lists that agree while the thing BETWEEN them does not exist — and this is the same failure one step downstream: the outputs were not wrong, the report was not wrong, and the missing thing was the artefact a person actually feeds to the next system. Name the file to use, in the output, first — and prove the naming by running the real producer into the real consumer. When a destructive REPLACE sits at the far end (a list rebuilt from whatever it was handed), also ask what the consumer does with an EMPTY input: here "held nobody" and "resolved everybody" are the same bytes, and only the file you upload decides which one the database believes.
A repair created the rows and left the old ones open — the CTE described a PRECONDITION and was reused as a POSTCONDITION
Symptom: house0197 reported students_created: 10 and looked like a clean run. The verification then showed held OPEN 165 and held resolved 0 — the ten students existed, but the ten held rows they came from were still marked waiting. For a few minutes ten people were simultaneously a placed นักศึกษา and listed in ยังนำเข้าไม่ได้, which is a state the app has no idea what to do with.
Cause: the script ran two statements against the same CTE:
sql
WITH ... only_team AS (... AND NOT EXISTS (SELECT 1 FROM students s WHERE s.person_id = p.id)),
pairs AS (...)
INSERT INTO students ... -- statement 1: gives all ten a students row
WITH ... pairs AS (...) -- statement 2: RE-DERIVES the same CTE
UPDATE student_import_unresolved ... FROM pairsonly_team is defined as "has a ทีม SAMO placement and no students row". That was true when the plan was made and false the instant statement 1 committed — the INSERT is what made it false. So pairs was empty, the UPDATE matched zero rows, and UPDATE ... FROM <empty> is not an error. The script printed a success count from statement 1 and said nothing about statement 2.
Fix: close the held rows by the fact that is still true afterwards — a held row whose รหัส and ชื่อ now belong to a real student — and assert the invariant at the end rather than trusting the write:
sql
select count(*) from students s join student_import_unresolved u
on <same รหัส> and <same ชื่อ> where u.resolved_at is null; -- must be 0Where it lives now: tools/house0197-promote-teamsamo-held.mjs, both the corrected UPDATE and the post-run assertion.
Rules:
- A CTE that describes a PRECONDITION cannot be reused as a POSTCONDITION. "Rows that still need X" stops matching the moment you do X. In a multi-statement repair, the second statement must key on something the first did not change — or better, on the thing the first statement produced.
UPDATE … FROM <empty set>succeeds. So doesDELETEandINSERT … SELECT. A repair that reports only what it inserted has not told you what it failed to update; print a count per statement and compare it to what you predicted, per statement.- State the end invariant and check it, not the individual writes. "Nobody is both a student and still held" is one query, survives any refactor of how the rows got there, and is the thing a reader of the ยังนำเข้าไม่ได้ tab actually depends on.
The year-admin CSV's ชื่อเล่น column was blank for every held row — the SQL just never selected it
Symptom. tools/house-year-sheets.mjs generates one CSV per รุ่น for a year admin to review, and the whole point of sending it to a year admin specifically (docs/HOUSE-YEAR-HANDOVER.md §a) is that they can recognise their own students "by หน้าตาและชื่อเล่น" — but the ชื่อเล่น column would have come out empty for every held (ค้างนำเข้า) row, silently.
Cause. The CLI's own SQL for the held population selected only student_id, first_name_th, last_name_th, sai_code as sai, cohort_year, resolved_at — no nickname_imported. toRow() calls the shared nickOf accessor (FIELDS.find(f => f.key === 'nickname').get, from census.js), which reads r.nickname || r.nickname_self || r.nickname_imported; held rows only ever populate the third one (students gets the first two from self-service, held rows never sign in). With the column absent from the query, nickOf(heldRow) was undefined for every row, always — not a bug a fixture test can catch, because the transform (buildYearSheets/toCsv) is tested against hand-built row objects that already have the field; the gap was entirely in the SQL string, which nothing in this repo runs against a real schema until a human executes it live.
Fix: add nickname_imported to the held SELECT (tools/house-year-sheets.mjs), and add a fixture test (tools/house-year-sheets.test.js, "a held row's ชื่อเล่น comes from nickname_imported") that at least pins the ACCESSOR side of the contract — this cannot catch a future SQL omission by itself, since the fixture supplies the field directly; it only proves toRow reads the right property name.
Where it lives now: tools/house-year-sheets.mjs, tools/house-year-sheets.test.js.
Rule: a hand-written SQL column list and the accessor that reads its result are two authors of one contract, and a fixture test only ever proves the second author, not the first — because the fixture is the thing a real row is supposed to look like, not the thing the query is supposed to produce. When a tool builds its own row shape by column-listing (rather than select *, itself avoided on purpose elsewhere in this repo — see the class-6 "a create table inherits too" entry), grep the list against every accessor the row is later passed through, not just against the columns the primary classification logic needs.
Revision pass, later the same night. The fix above shipped disclosing its own limit — "cannot catch a future SQL omission by itself" — as a TODO rather than closing it. Reintroduced the exact original bug (dropped nickname_imported from the held SELECT again) and reran the full suite: 14/14 still green, confirming the fixture test really cannot see a regression in the SQL string, exactly as claimed. Added a second, source-text test (tools/house-year-sheets.test.js, "main()'s held-population SELECT names nickname_imported") that reads the tool's own file text and asserts the column appears in the held query specifically, not just anywhere in the file. Reran the same reintroduced bug against the new test: red, as expected. This is a class-7 "source guard is a review, not a test" — it does not run the SQL, so a typo'd column name it would still miss — but it is strictly more than nothing, which is what the fixture test alone provided against this exact regression.
"Run failed: build" on a push whose npm test was green — the suite is not the same suite on both machines
Symptom (as reported): "why the build workflow run keep getting error like this. you should improve your workflow or test or something to detect so that this wouldn't occur." Two pushes went red on CI (ef1fe7a, bee6d90) while npm test was 2,176/2,176 on the laptop that made them.
Cause: state-handoff.test.js sweeps every path STATE.md and docs/state/*.md name and fails on one that is not there. It asked existsSync. externaldata/ holds the raw handover files — 1,776 real students — and is gitignored, so it EXISTS on the maintainer's laptop and does NOT exist on a CI checkout. The sweep therefore answered a different question on each machine.
Both directions happened on the same night, which is what makes it a shape rather than a slip:
| assertion | laptop | CI |
|---|---|---|
| the two dead-pointer sweeps | green | RED |
| "no exemption survives the file arriving" | RED | green |
The second was fixed first, with a local git check-ignore helper — and the two sweeps ten lines above it were left reading the filesystem. One rule, two implementations, one fixed: .claude/rules/mistakes.md class 6, inside the file whose job is to catch exactly that.
Fix: one missingFromRepo() asking git, used by all three sweeps. A gitignored path is not a repo file — it is a local artifact a note may legitimately name — so it is skipped whether or not this machine has it. The two externaldata/ entries in ABSENT_ON_PURPOSE then became unreachable and were deleted: an exemption that is never reached silences that path for ever, including after a real rename.
And the part that catches the NEXT one: npm run test:clean (tools/test-clean.mjs) stages exactly git ls-files into a temp directory and runs the suite there. Demonstrated on the restored bug — npm test 15/15 green, npm run test:clean 1 failed, which is what CI would have said.
Two things it needs, both learned by getting them wrong:
- It must be a real git repo (
git init+ the same remote), or every assertion that shells out to git goes red for "not a repository" — a clean-room that reports failures CI will not report teaches you to ignore it. - It must see the HISTORY (
.git/objects/info/alternatespointing at the real object store), orgit cat-file -e <sha>says no for the perfectly good sha in the ✅ DEPLOYED line. CI checks out withfetch-depth: 0and resolves it fine.
Where it lives now: src/js/state-handoff.test.js (missingFromRepo), tools/test-clean.mjs, npm run test:clean.
Rules:
- AN ASSERTION THAT ASKS THE FILESYSTEM ASKS A DIFFERENT QUESTION ON EACH MACHINE. Anything gitignored —
externaldata/,.env.local,dist/— is present for the author and absent for CI. Ask git what the repo contains;existsSyncanswers what this disk contains, and those are not the same set. - A green
npm testis not evidence about a push. If the suite reads anything outside git, run it where only git's files exist before believing it. - When you fix one reader of a rule, grep for the others in the same file. The fixed one and the broken one were ten lines apart.
The night agent's memory lived on the VM and nothing ever synced it
Symptom (as reported): "you should also sync memory system to the vm all the time, or the claude night agent would work wrong". Checked: the VM held 51 files dated Sep 11; the laptop held 55.
Cause: the memory directory is deliberately OUTSIDE the repo — some entries name real students and this repo is public, so git add -A must never reach them. That is right, and it means nothing syncs it. install.sh copied the scripts and the systemd units and not the memory, so the VM's copy was whatever the last person had hand-copied, eight days earlier.
Why it is worse than a missing file. The four entries the VM did not have were the ones written to say that earlier numbers had CHANGED — "13 held rows with no รุ่น" is now 0, "18 สาย problems" is now 5. An unattended run would have read the superseded figures as current, from a file that exists and reads fluently, and acted on them. Nothing in the night's output would have looked odd. It is the same shape as STATE.md's retyped deployed sha: the stale copy was the instrument.
Fix: install.sh now rsyncs the memory on EVERY install — not behind a flag, because a sync you have to remember is a sync that is sometimes skipped and the failure is silent — with --delete, so a memory removed for being WRONG disappears there too. It refuses to arm at all if the directory is missing, rather than running an agent that reads nothing.
And because an installer only helps when someone runs it, run-night.sh now prints the memory's file count and age into the Discord message a human actually reads, with a warning past three days. A stale memory cannot be noticed by reading it; the only tell is its date.
Also fixed in the same pass: install.sh used to systemctl enable --now every time, so syncing the memory would have silently re-armed a timer the owner had deliberately disabled. Arming is now --arm, typed on purpose.
Where it lives now: server/night-agent/install.sh, server/night-agent/run-night.sh (the MEM_LINE block), skills/night-agent.md.
Rules:
- STATE THAT LIVES OUTSIDE THE REPO HAS NO SYNC UNLESS YOU WRITE ONE, and its staleness is invisible by construction — there is no diff to see and no build to fail. Name the thing that copies it, and make it unconditional.
- A remote copy of anything should report its own age where the operator looks. Not in a log nobody opens: in the message that is already being sent.
- Deploying is not enabling. A script that installs and arms in one step will eventually undo a deliberate "off".
shop0202's "non-Google slip URL is refused" was GREEN on the code it exists to catch
Symptom: running the new proof against production's OLD function bodies, 13 of 14 DENY rows went red, and one stayed green: "self: a non-Google slip URL is refused".
Cause: the evil URL was a hand-escaped JSON literal inside a JS template literal inside SQL. One escaping layer ate a backslash, the ::jsonb cast failed, and the exception when others branch recorded "refused". It was refused, but for the wrong reason.
Fix: the value is built with jsonb_build_object(… chr(34) …), so there is no hand escaping. Re-run: 14/14 red on the old bodies, 22/22 green on 0202.
The general rule: the ritual caught it — reintroduce the bug and watch EVERY deny go red. A catch-all handler scores any error as the refusal you wanted. Where the expected refusal has a name, assert the name, not "an error".