The whole product rests on one idea: a model checking its own work is not a check. So the checking is done by a different model, with its own hands, and what it did is shown to you.
The verdict chip
Every reviewer's turn carries a chip. It says whether the seat agreed, and it is coloured by the strength of the check behind it, not by the conclusion. Open it and you see the commands that seat ran and whether they passed.
A reviewer that agreed without opening a file or running anything is labelled as exactly that. That is a weaker review than one that checked, and the app would rather tell you than flatter itself.
The Summary
At the end of a run, one block: the outcome, how many reviewers agreed, how many commands ran and how many passed, what files changed, and what was found. Read it before the answer — it tells you how much to trust what you are about to read.
Findings
A review's deliverable is what it found. Defects named anywhere in the run are gathered into a list rather than left scattered through the prose, where a genuine bug can be mentioned once, agreed to be out of scope, and quietly lost.
Blind peer review
While it is on, every room — the Mac, the web and the phone — is anonymous: the seats read each other as "Model A", "Model B" and rank the others without knowing who is who, and the person always reads the real names. The ranking is asked for once, after every seat has answered, so every answer is ranked by every other seat; asked during the turns instead, a seat could only order what was written before it spoke, and whoever answered last was read by nobody and missing from the result. Each seat sees its own answer nowhere in the list, so it has nothing to favour. The transcript shows the real names; the receipt shows the ranking. It is the closest thing to evidence available when nothing can be executed, and it stops a seat deferring to a brand. It costs one short extra question per seat per round, and "Have the AIs rank each other" in the room settings turns it off.
Replay checks
When a run changed files, the app re-runs the checks the room established, after everyone agreed. A replay that fails outranks the room's conclusion: the answer was wrong, found by the machine rather than by you a day later.
The audit log
Every message in a discussion is hash-chained, and the discussion carries a digest an auditor can recompute. If you need to prove a transcript was not edited after the fact, that is how.