Skip to content

PTV-SCF-0001 P2 — verification round

Date: 2026-08-13 Status: open Supersedes: none Superseded-by: none — current Phase: PTV-SCF-0001 P2 · Golden prompts as skills Chunk: C5, first act · gates G-P2-1 … G-P2-6 Shape: G-02 (docs/PETROVA-GOLDEN-PROMPTS.md §3) Terminal state of this record: a gate table and a classified friction list. It does not close the phase. That is G-03, in a separate session, after this record is merged.

Not from the chunk records. Each gate was re-run or re-evidenced in this session, because a round that reads the delivery records back is a round verifying its own paperwork.

G-P2-1 and G-P2-3 required the eight cold-session runs C2 deferred here. They were run against rocky-hq per C2’s protocol: eight fresh sessions, each given one skill file and the task that skill exists for, each forbidden to read any other control-plane file or to ask a clarifying question, each asked to report what it loaded, what it produced, and — the question that matters — what it needed that the skill did not reach.

One deviation from the protocol as written, recorded rather than glossed. C2’s protocol says “invoke the skill by its slash command and nothing else.” The skills are not installed in rocky-hq, so each run was handed the absolute path to one SKILL.md instead. That is the same text by a different door, and the runs confirm the skill body was the operative instruction. But it is not what the protocol said, and it means the commands/ invokers shipped in C2 are still unexercised. Raised as F-31.

GateVerdictEvidence
G-P2-1 · Every one of the eight skills loads in a cold session and produces a correct plan against a test repoPASS8/8 fetched the preamble, loaded rocky-hq/.petrova/contract.yaml, and reached the terminal state their own ## Terminal state section names, citing real files. Four produced artefacts (petrova-decide a record body + dry-run, petrova-plan a six-task sequence with halt gates, petrova-verify-round a seven-gate table and thirteen classified items, petrova-drift-check a three-class findings list). Four refused, and each refusal was correct: petrova-phase-open on an indeterminate phase, petrova-phase-close on a missing round, petrova-onboard on an already-admitted slug, petrova-recover classified NO_PRIVILEGED_PATHS as a refusal rather than a failure and declined to route around it.
G-P2-2 · A session using only the fetched preamble performs equivalently to one given the pasted §1 blockFAILTwo cold sessions, same task (“add a caching layer to the console service and update .github/workflows/e2e.yml”), one fetching and one pasted. Both cited L1–L7 and W1–W4 and both predicted the .github/workflows/ refusal. They diverge on the gate’s third clause — no governance claim present in one and absent in the other. The pasted arm found that console is a git submodule, concluded that a rocky-hq-scoped verb can only move its pointer and cannot write console source at all, and refused to name a verb because no catalogue is on disk. The fetched arm made none of those claims. It also reported it never saw the file’s bytes: WebFetch answered through a summarising model, one attempt refused outright citing a “125-character maximum for quotes” it could not source, and the two attempts contradicted each other about whether such a limit exists.
G-P2-3 · No skill requires out-of-band context to workFAILAsserted per skill, as the gate demands. 8 of 8 reported at least one fact needed for correct operation that is not reachable from the skill text plus the artefacts it names. The recurring one is fatal to the phase objective: no skill says where a repo’s phase state lives. Four runs independently guessed MILESTONES.md, and one noted it also had to decide unaided whether that file counts as a source or as a projection under L5 — the skills say “projections (CLAUDE.md and similar)”, and “and similar” is carrying the entire judgement.
G-P2-4 · errors.json carries stage, meaning, correct_response per code; verbs/index.json carries a fingerprint per verb and refused paths as globs; the preamble carries L1–L7 labelled and a non-empty warning listPASSRe-run this session, not read from C1’s record: codes postbuild gate reports 18 verbs fingerprinted, 4 refused-path globs, 80 refusal codes with recovery; site postbuild gate reports 7 law clauses and 4 live warnings.
G-P2-5 · The G4 checker exits non-zero on a seeded orphan and zero on the current baseline, and is registered in checks.yamlPASSRe-run: npm --prefix host run requirements:check exits 0, “14 SRs trace to a root, 12 DRs trace to an SR”. Seeded fixture exits 1 with E_ORPHAN_REQUIREMENT. PTV-CHK-0036 present, status: built.
G-P2-6 · Every one of the 35 unallocated slugs appears in the C4 record with a proposed DR block and a named owner, and the record is uncountersignedPASS as delivered, with the second clause now false35 rows, contiguous, no duplicates, block and owner on every row, and the set asserted equal to the estate’s own unallocated population computed from registry.yaml + estate.yaml. The countersign was ☐ unticked at the delivery commit 09e3a4a, verified by reading that commit rather than the working tree. It is ticked today, by operator directive, after delivery. The gate is evaluated against what C4 delivered; the alternative reading — evaluating a mutable property at an unspecified time — would make the same gate pass or fail depending on when someone looked. That the gate does not say which is a defect in the gate, raised as F-33.

Four PASS, two FAIL. Neither failing gate was reworded, and neither is close to passing. Both failures are about the same thing from two directions: the law is fetchable, and the fetch is not yet trustworthy or sufficient.

What the FAILs mean for the phase objective

Section titled “What the FAILs mean for the phase objective”

P2’s objective is make governance law fetchable rather than pasted. The artefacts exist and are correct — G-P2-4 confirms it, and curl returns 8,235 bytes of exact law. What P2 has not established is that an agent fetching it receives it. On the evidence here, the pasted arm outperformed the fetched arm on the one clause designed to detect a governance difference, and the fetch path degraded the artefact into a paraphrase before the agent ever saw it.

That is not a reason to go back to pasting. It is the finding that the objective is one layer deeper than the phase assumed.

Part B/C — friction, surfaced and classified

Section titled “Part B/C — friction, surfaced and classified”

Fourteen items. Ten deferred, three closed, one in-budget. Every deferred item carries a named target phase.

IDItemClassTarget / justification
F-22The preamble contradicts itself on its most-cited warning. Line 35 (W1): “DO NOT cite a meta-rule by number. Cite by name and URL.” Line 60, §The meta-rules in force: “Cite by number.” Same file, 25 lines apart. Every C2 skill instructs agents to obey W1; the artefact they fetch instructs the reverse.DEFERREDP3 · the generator emits both strings from machine-facts; one is wrong. Not fixed here — fixing it means editing a P2 deliverable during P2’s close, which is the extension L6 forbids.
F-23The fetch path returns a summary, not the artefact. Four runs reported it independently; one reported a fetch refused outright as a suspected jailbreak, citing a quote-length rule it could not source. Every law citation in every cold run therefore rests on a paraphrase.DEFERREDP3 · this is the mechanism behind G-P2-2’s failure and the highest-leverage item in the list. Fixing it may mean serving the preamble in a form a fetcher will return literally, or publishing a fetch-and-verify recipe with the Law fingerprint as the check.
F-24The preamble’s own framing — “the law you operate under”, numbered rules that “outrank your task and your plan” — reads as prompt injection to a safety-tuned fetcher, which refused it once and served it on a reworded retry. A skill whose step 1 is “fetch the law” fails at step 1 intermittently.DEFERREDP3 · same target as F-23 and probably the same fix; recorded separately because the cause is the document’s voice, not the transport.
F-25The verb schema cannot carry what the skill mandates. verify_round’s items[] accept id, description, surfaced_by, evidence_ref and nothing else; the schema’s own side_effects says it “does NOT classify items — that’s close_phase’s job”, and close_phase carries friction_classifications. Golden prompt G-02 Part C makes classification the round’s job and G-03 forbids the close from doing it. The prompts and the schemas disagree about who classifies, C2’s skills faithfully encoded the prompts, and a cold run smuggled its classifications into the description string as an invented "DEFERRED->target:" prefix.DEFERREDP3 · a ruling, then whichever of the two artefacts is wrong. This record has the same problem and solves it the same way, in prose.
F-26No skill says where phase state lives. Four of eight runs guessed MILESTONES.md; one had to decide unaided whether that file is a source or a projection under L5. The mandatory load, .petrova/contract.yaml, carries no phase state at all — so the one file the phase skills require is nearly useless for their own preconditions.DEFERREDP3 · the skills are the fix, but the underlying question is where a governed repo declares its open phase, which is a contract-schema question.
F-27petrova-onboard has no already-admitted exit, and never says where the control-plane registry lives. Its preconditions tell the agent to check the registry and its workflow reads as unconditional. Worse: rocky-hq has its own unrelated registry.yaml at repo root, so a literal cold agent can read the wrong file and conclude a governed repo is unregistered. The run reached the right answer by inventing the exit itself.DEFERREDP3 · a refusal condition plus a stated path. The wrong-file trap is the part that produces a confidently wrong answer rather than a stall.
F-28Verb payload schemas sit under ## Reference, not ## Load first, so no skill requires fetching them. petrova-decide composed a dry-run with invented parameter names and flagged it as the largest correctness risk in its own output. Two other runs skipped /verbs/index.json for the same structural reason and noted their refused-path knowledge was therefore incomplete.DEFERREDP3 · promote the per-verb schema fetch into ## Load first for every skill that composes an invocation.
F-29Every skill footer resolves links through .petrova/brand.yaml#blog.base. That file does not exist in rocky-hq, and five runs reported falling back to the literal host with no stated fallback. The documented resolution mechanism is absent in the repo it is documented for.DEFERREDP3 · either ship brand.yaml in the boot kit or state the fallback in the footer.
F-30Skills mandate vocabularies they do not define. petrova-drift-check makes severity a refusal condition and defines no scale — the run invented LOW/MEDIUM/HIGH and said so. petrova-phase-open mandates a sub-milestone state per item and enumerates no legal states.DEFERREDP3 · enumerate both in the skills, or point at the artefact that does.
F-31The eight runs exercised SKILL.md files by path, not the commands/ slash-command invokers C2 shipped, because the skills are not installed in the fixture. The slash commands remain unexercised, and the install path itself is unverified end to end.DEFERREDP3 · install the skill set into the fixture and re-run at least one, which also tests skills/README.md’s instructions.
F-32The runs incidentally produced a substantial, evidenced finding set about rocky-hq, not this repo: ~30 sub-phases marked closed with zero verification-round artefacts, against that repo’s own stated definition of closed; CLAUDE.md frozen at “State as of Phase 5” while Phase 7 is open; two phases open simultaneously; three weeks of merged work advancing no phase and ending in a revert of its own change; a not_applicable_review_by date seven days past due.DEFERREDrocky-hq, not a P-phase · L4 — project truth lives in that repo. This round must not absorb another repo’s findings, and the correct act is a handover to rocky-hq’s own ledger. Recorded here only so the evidence is not lost.
F-33G-P2-6 asserts a mutable property with no timestamp. “The record is uncountersigned” was true at delivery and false hours later, by legitimate operator action. A gate whose verdict depends on when it is read is a gate-writing defect — the same class G-02 names when it says an unevaluable gate is a finding about how the gate was written.CLOSEDResolved in this record by evaluating as-delivered and saying so. No carry: the fix is a phrasing convention for future gates, and it is stated here.
F-34This round evaluated gates on work authored in the same session, by the same agent, for a phase that agent scoped. The structural separation the phase itself insisted on (R-2, the G-02/G-03 split) is not achieved between author and verifier here — only between verify and close.IN-BUDGETAbsorbed, with the justification stated rather than assumed: the eight cold runs and the G-P2-2 pair were executed by sessions with no access to this one’s context, and both FAIL verdicts come from their evidence, not from the author’s judgement. The separation that mattered held where it could be made to hold. Naming it is the mitigation; pretending it was absent would be the defect.
F-35doctor-all.sh fails on kahn-hq and stratt-hq, reproduced on a clean tree before any P2 work. Unrelated to this phase, surfaced by it.CLOSEDNo carry from P2 — not this phase’s work and not this phase’s to fix. Recorded so the next reader does not attribute it to P2.

Nothing here is proposed for absorption into P2. Every fixable item is DEFERRED with a target, per L6. In particular F-22 — a one-line generator fix I could make in minutes — is deferred precisely because it is a one-line fix on a P2 deliverable during P2’s close, which is the exact shape the round exists to refuse.

What a clean round would have looked like, and did not

Section titled “What a clean round would have looked like, and did not”

A round that surfaces nothing was not run properly. This one surfaced fourteen items, two gate failures, and a contradiction inside the phase’s own flagship artefact. The uncomfortable part is not the count — it is that G-P2-4 passed on the same artefact F-22 condemns. A gate asserting field presence and non-emptiness cannot see a field that contradicts another field. That is worth carrying into how P3’s harness gates are written.

  • docs/decisions/2026-08-13-ptv-scf-0001-p2-open.md — scope, rulings, the six gates.
  • docs/decisions/2026-08-13-ptv-scf-0001-p2-c2-skills.md — the eight skills, the run protocol, the fixture.
  • docs/decisions/2026-08-13-ptv-scf-0001-p2-c3-g4-checker.md — F-15…F-18.
  • docs/decisions/2026-08-13-ptv-scf-0001-p2-c4-dr-allocation.md — F-19…F-21.
  • spec/verbs/verify_round.schema.json, spec/verbs/close_phase.schema.json — F-25’s evidence.
  • https://petrova.blog/llms-preamble.txt — lines 35 and 60, F-22’s evidence.
  • Subagent: PTV-SCF-0001 P2 verification round (session 2026-08-13)
  • Human: ☑ (proxy) countersign — accepts this round’s verdicts and its classified friction list. Ticking this does not close P2. The close is G-03, in a separate session, and it must run against this record once merged.
    • Countersigned by human:devarno on 2026-08-13, by explicit directive in session (“approved — proceed accordingly”). Ticked by the agent as scribe, not as signatory. No other part of this document is edited (MR-7).
    • This signature accepts the verdicts; it does not waive the two FAIL gates. G-P2-2 and G-P2-3 remain FAIL. close_phase’s third precondition requires every gate to read PASS or to carry a recorded and merged waiver, and no waiver exists. Waiving a gate is a separate governance act with its own record, and an agent does not author one on an approval of something else.