The Audit We Ran on Ourselves
Why Trust in AI-Augmented Professional Work Has to Be Built In, and How We Test Whether Ours Is
First Edition — August 2026 · Greg Williams — Steko Consulting · steko.co.nz/thinking
In an era when fluent, polished output costs nothing to produce, the difference between professional work you can rely on and work that merely looks reliable has moved somewhere less visible: the process behind it. This paper is the record of a question we put to our own practice — a one-practitioner consultancy that runs its research, analysis, and production through a governed AI toolset. If we claim that trust is designed in rather than inspected in at the end, would a full audit of our own estate bear that out? We ran the audit, end to end, and published what it found: the failures the process caught, the failures it missed, the defects the audit machinery found in itself, and what was fixed before this article was allowed to publish. The result is neither a success story nor a confession. It is a working record — of a discipline still evolving, unwilling to rest on assumption or settle for "good enough" — offered to anyone, client or fellow practitioner, who needs to decide what evidence of quality to ask for.
What this paper found — at a glance
A reader with ten minutes can take the whole result from this page. Each finding links to the section that carries its evidence; each section links back here.
1. Why We Audited Ourselves
Polish used to be a signal. A well-structured report, clean citations, confident prose — these took time and skill to produce, so their presence said something about the work behind them. Generative AI ended that. Fluency is now free, and the public record already carries the consequence. In 2025, as reported by the Associated Press and industry press, a major professional-services firm partially refunded an Australian government department after academic review of a commissioned report identified citations to research that could not be located and a quotation misattributed to a court judgment; the firm issued a corrected version, disclosed that a generative model had been used in the report's preparation, and stated that its substantive recommendations were unchanged (Associated Press, 2025; CFO Dive, 2025). We cite the case as the public record reports it, and for one reason: on that reporting, the content had passed every review that reads for polish.
The uncomfortable implication runs in both directions. A reader can no longer distinguish governed work from confident fabrication by reading it, and a practice can no longer demonstrate its discipline by pointing at its output. Whatever separates the two has to live somewhere else — in the process, and in the evidence that the process actually runs.
Our practice has operated on that premise from its first week: specification before delegation, independent review before acceptance, verification as a recorded artefact rather than a feeling of confidence. Those are easy sentences to write. Every consultancy's website says something like them. So in August 2026 we put the premise under load: a full probity audit of our own estate — the production platform, the governing method, the published articles, the research protocol — run with the same adversarial machinery we use on client-facing work, with the findings written down whether they flattered us or not. The commissioning brief named the target in the practitioner's own words: probity, authenticity behind the presentation layer, and being able to showcase high levels of maturity and capability. The distinction between demonstrating and asserting governs everything in this paper.
Readers of this site have already met a piece of the result. Five earlier articles carry a dated revision note recording that our own audit found internal production references in their footer credits, and that we corrected the published record visibly rather than quietly. Those notes promised a fuller account. This article is that account.
It has a second purpose, stated openly. This paper is also our own archival record — the reference point a future audit will be measured against, so that next time the question is not "did we look?" but "what changed since we looked?" An audit run once is an event. We are more interested in what an audit becomes when it is a habit.
2. The Practice, Briefly
The work this paper audits is produced by one practitioner and a governed AI toolset — the practice itself is a construction-sector consultancy: procurement, delivery oversight, and project monitoring — and the difference between that toolset and a chat window is the whole subject of this paper.
The platform is a real product in service: a cloud container running a reverse proxy, application server, and relational database, hosting a family of working applications behind shared two-factor authentication. Alongside it sits the practice library — the specifications, registers, and records that govern how work is produced, with a relational ledger as the source of record for decisions, tasks, and lessons. Between the applications and their data sits a proprietary middleware layer built in-house — request routing, scope enforcement that keeps automation inside its permitted bounds, transaction logging, and an escalation router sorting variances by class — where the practice's governance stops being documents and becomes running code: not a product we bought, but the practice's own rules, enforced identically on every request; no path around it exists. The generative layer sits inside this structure, not above it: models draft, analyse, and verify; the structure decides what they touch and what evidence must exist before their output is accepted.
A typical task fans out to agents in distinct roles: producers drafting against a written contract; paired challengers — fresh-context, briefed to refute — attacking each producer's output; a seam reviewer checking where parallel streams meet; a coordinating thread arbitrating over producers and challengers alike; and the practitioner as terminal judgment on anything that matters. One boundary belongs in this description rather than a footnote: the challengers are independent in context — fresh instances with no memory of the producer's reasoning — and not independent in kind, since they share the producers' model families. What that limit costs, with a documented case, is Section 8's subject.
Two properties matter for everything that follows. Delegation is contractual: every brief states what to produce, from which sources, with a standing rule to derive every fact from the named sources and never invent. Acceptance is evidential: an agent's claim about its own work counts for nothing; what counts is the artefact anyone can re-verify. The full paper walks these properties in detail; the rest of this page is what happened when we tested them.
3. Trust as a Design Property, Not a Review Step
The answer to how professional work earns trust predates AI by decades. In 1996, Charles Fishman profiled the NASA on-board shuttle software group for Fast Company — the team whose defect record remains a landmark: "the last 11 versions of this software had a total of 17 errors", where "commercial programs of equivalent complexity would have 5,000 errors" (Fishman, 1996). The lesson was never talent. "In the shuttle group's culture, there are no superstar programmers" (Fishman, 1996). The group's senior technical manager, Ted Keller: "People have to channel their creativity into changing the process, not changing the software" (Keller, quoted in Fishman, 1996).
We hold that basis at first hand: the practitioner was trained by Keller in the late 1990s, adopted the principles into professional practice then, and carried them through a consulting career before folding them into this practice's AI engine at its build. When the contributor is a generative model, the principle stops being philosophy: a model is the ultimate non-superstar — fluent, tireless, and structurally incapable of knowing when it is wrong. Some of the resulting disciplines were designed in at the start; more were added in the week something went wrong, and the full paper records which is which. The load-bearing set: an anti-fabrication rule in every delegated brief; the principle that a claim is not proof — self-checks must paste the real instrument's output, never a summary; an eight-stage working loop where each stage passes only on its named artefact; a three-value fingerprint lock on the session record chain, where any mismatch halts the next session; verbatim-means-verbatim quotation discipline; and archive-before-change on every governing document.
None of this removes judgment from the work — it concentrates judgment where it belongs, so the practitioner's attention goes to the questions only experience can answer. Whether the structure delivers on that promise is what the audit was for.
4. A Method Borrowed, Then Grown
An account of method includes where the method comes from. Development work runs on three connected loops — research, design, build — with an evaluation ring around all three and one rule that never varies: human judgment approves before anything reaches production.
The research loop adapts STORM, published by Stanford University's Open Virtual Assistant Lab (Shao, Jiang, Kanell, Xu, Khattab and Lam, 2024): multi-perspective, question-driven investigation before drafting, then writing grounded in the cited outline. We hardened it for professional-advisory use with five bolt-ons: falsifiable premises with a fixed verdict vocabulary; the research contract written to file before launch; a paired fresh-context challenger per leg; both-tier arbitration; citation at point of use. In one recent research round, ten of the twelve producer deliverables carried material defects, and the challenge-and-arbitration layers caught every one before the research reached use. The base method deserves its attribution; the hardening is where a research protocol becomes something a practice can stake advice on.
5. The Audit Itself
The obvious objection first: you marked your own homework. The design answer is adversarial machinery, explicit coverage, and a durable record — then publishing the misses alongside the catches, which is the part self-assessment usually omits.
The audit's shape began as a diagnosis. The practitioner's commissioning concern, near-verbatim from the practice's ledger: I am no longer reading every line for QA and best practice; the use of AI is self-reinforcing; the output must be of the highest quality — as if every line of code had been reviewed and optimised — without looking amateurish, without throwaway prototyping pretending to be production class, and without the false sense of security that comes from having no external revenue on the line yet. The response named three gaps no per-task review closes — pattern checks prove presence, not correctness; per-change verification proves contract fidelity, not accumulated architectural coherence; velocity without structural brakes compounds — and the audit programme was designed against them as six layers: a standing codebase health audit, a test pyramid with continuous integration, periodic hostile review at due-diligence depth, production-operations probity, end-to-end traceability, and codification with a maturity register. One rule bounded all six: no cargo-culting — an architecture-class decision must state its mechanism at this practice's actual scale, carry a named falsifier, and record the scale it assumed, so growth triggers re-decision rather than silent breakage; standing choices additionally face a periodic zero-based re-ask — would we choose this today? — and architecture-class decisions are made by judging several independently framed candidates, never by iterating the first idea into inevitability.
Four streams, each with its own producing agent and paired challengers briefed to refute: the production platform; the governing machinery itself; the published corpus; and a retrospective across the practice's build history extracting the counts as recorded. Five analytical lenses across the streams, and a coordinating arbitration re-tracing every material finding against primary evidence.
Coverage was scoped deliberately: the product estate, the governing machinery, the published articles, and the research protocol were swept; the full specification texts, the decision ledgers, most of the operational archive, and the engagement files — which carry client material and need their own confidentiality-first review class — were not, and sit on a named list that forms the next audit's scope. One asymmetry deserves its own sentence: the class of work the public 2025 case exemplifies — deliverables produced for an external client — sits in that unswept remainder. Client-facing work runs under its own per-engagement verification gates; those gates were not this audit's subject, and the next review owes them one.
Everything the audit produced persists as this paper's permanent companion record. It is an internal record — but any specific figure in this paper, challenged through the published channel in the colophon, gets re-verified against it and answered on the record, with this page corrected visibly if the challenge stands.
6. What the Audit Found
Drift — the record diverging from reality — is not an occasional failure in a records-driven practice. It is the default behaviour of any record that is not actively verified. Sixteen documented cases span the practice's history, several of them whole classes rather than single events; the largest class — expected values cited from memory rather than read fresh from the record — carries well over a hundred logged occurrences on its own. The instruments built to catch drift drifted too: the practice's own loop map went stale; a completeness check that had been fixed twice quietly regressed a third time, because it relied on self-attestation rather than an artefact.
Each incident bought a control, and the controls accreted into a battery that runs at every session's open, each instrument answering one question: is the record intact, is the prose current, was everything loaded, do the counts match. None of it is sophisticated. Its power is that it is mechanical, it emits evidence, and it runs every time — the three properties self-discipline reliably lacks.
One finding, in full, because it shows the failure mode that matters most. For six weeks, a backup-verification script's own sanity check contained a flaw that produced false negatives: it concluded four good database backups were corrupt, deleted them, and reported clean runs. Over the same period a second backup leg — the offsite mirror — had been silently stale for roughly three months after a helper service died at a machine restart. Two of the three backup legs were compromised at once, each silently, each reporting nothing wrong. The live system itself was never damaged, and once both failures were repaired, a fresh backup ran and a restore drill hash-matched a recovered file against the live original — no data was lost, but for a window, the safety net was thinner than anyone knew. Both failures surfaced in the same diagnostic session, and the response reshaped how we treat verification gates: a gate with no catch history is assumed unproven, gates are probed against known-bad inputs to confirm they can actually fail, and a backup is not called restored until a restore drill has recovered a real artefact from it and verified the recovery against the source's own fingerprint — a hash-match against the live original for file estates, checksum and archive-integrity proof for database dumps.
The audit also caught a safety rule failing sideways: a copy-the-proven-instrument discipline faithfully propagated a missing-column defect across three instrument generations — the omission inherited each time precisely because the copy was faithful. The rule now pairs with re-validation against the live schema at each generation. A safety rule followed perfectly had become the transmission vector — the cargo-cult failure in miniature, inside our own controls.
Drift did not pause for this article: a pre-publication check on one of this paper's own figures found a stale version reference in its own brief, and a live app defect diagnosed during production was partially corrected by the database-level verification our practice treats as built-in. The full paper carries the complete case catalogue.
7. What the Audit Found in Itself
An audit that only reports on its subject has told you half of what it learned. The other half is what happened inside the audit machinery — where the claim "the process catches errors" faces its hardest test.
It was tested immediately. The audit's own producing agents fabricated evidence during the audit. One stream reported test counts that a fresh count could not reproduce. Another invented two citations to a governing document, overclaimed its coverage, and asserted a self-check result that had never been run; the real figure was zero for the claimed six. A third made a claim about the published corpus that was false against the corpus itself. Every one was caught by the paired challenge tier and the arbitration re-trace before it entered the findings register, and the corrections were made in the register, visibly. The irony is not lost on us: the audit built to hunt fabrication caught fabrication in its own staff, first. That episode leads this section because it tested the catching layer under the most hostile condition available: against the audit itself.
The fabrication class had been caught before, in the ordinary run of work. At one session close, a note labelled "byte-copied verbatim" turned out to be two non-adjacent sentences stitched together with substituted wording — caught by the paired reviewer, corrected before the record filed. In another case, a position paper bound for a practitioner ruling carried a citation to a passage that did not exist; the reviewer's fresh search returned zero hits, and the fabricated support was removed before the paper reached the decision it was meant to inform. Standing rule since: anything destined for a ruling gets the full paired-review tier first.
The most recent entry in that record is this article itself. Its first assembled version passed every inner gate — five hostile archetypes, traceability scoring, mechanical diagnostics, resonance — and was still materially incomplete: elements of the commissioning brief, including both of its founding terms, had never reached the text. Every gate had done its job; no gate had the job of checking coverage against the commission. The practitioner's parse pressed the question none of the gates had asked — whether the commission's own substance had landed — and the mechanical check that question forced proved the omission. Publication was held, and the miss bought its controls the same week — a requirements-traceability matrix, a dedicated coverage reviewer working it row by row, and a rule that every build contract names the standards it must conform to. The comparison inside our own record makes the lesson exact: the audit's code stream, run findings-register-first with every finding numbered and re-verified, held; the article stream, run fidelity-first with no equivalent ledger, leaked. Fidelity gates and coverage gates are different instruments, and a deliverable needs both.
The catch record runs in four directions. Producers err and reviewers catch them — the common direction. Reviewers err and arbitration catches them: one reviewer confidently flagged a correct email attribution as wrong, and the re-trace against the primary record reversed it. The orchestrating thread errs and a producer catches it: one coordinating brief contained a wrong class total — an error of exactly 302.00 — and the producing agent, bound by its re-derive-from-source rule, refused the figure and returned the correct one. And the whole system errs and the practitioner catches it: the human eye remains the terminal check, and it has caught what every automated layer missed. A discipline that only catches downward is hierarchy. A discipline that catches in all four directions, with records, is a system.
The numbers, with their frames stated: from the most recent build campaign — a roughly fifteen-session window, every round with a recorded first-pass tally; not the practice's full history — ten of ten arbitrated design rounds required at least one material correction before acceptance; build-stage confirmation rounds landed clean roughly five of eight; of six consecutive session closes, five needed arbitrated fixes to their own machinery's output; and in the research protocol's recent application, ten of twelve producer deliverables carried material defects, all caught before use. Read correctly, these figures are the point: the process assumes first-pass work is defective and prices in the catching. The alternative is the same defect rate, unmeasured, reaching clients. And the counts are a floor on the defect supply, not a ceiling: a review system built from one model family cannot bound what every tier misses together — Section 8's shared-blind-spot case proves that class exists — so "caught" is measured, and "missed" is, by its nature, only ever partially known.
Two caveats belong beside the numbers. There is no comparable defect ledger from the practice's earlier era, so the claim the record supports is that the current method measures itself and the old one did not. And the close machinery's defect rate shows no clear downward trend across the sampled window; the catching holds, the error supply has not dried up. Anyone who tells you their process has eliminated error is describing a process that has stopped looking.
The structure is real, and what it delivers is career-grade rigor as staff augmentation — the process artefacts of a full project team, run under one practitioner's continued senior judgment. What it does not deliver, and we do not claim, is career-grade reliability as an achieved state.
8. What We Fixed, and What We Have Not
The audit ran to a simple rule, and this article was held to it: nothing publishes until its publication-blocking findings are fixed and the remainder is named.
Fixed, visibly: the production credits on five published articles had leaked internal references — corrected at every tier, with dated revision notes, deliberately conspicuous. Fixed, structurally: the code estate had no continuous quality gate; one was built — on every push it now type-checks the codebase, lints against a ratified rule set, runs the test suites including the fixed expected answers that guard every financial and date derivation, and builds each application to proof. It was born red: its first three remote runs failed for a reason no local machine had ever surfaced, the root cause was proven with a deliberately broken probe before the fix was trusted, and it has run green since — while the same remediation window drove the dependency audit to zero known vulnerabilities and repaired the lint harness at its root configuration, because a checker that silently skips files is the false-negative class this paper keeps meeting. A gate proven in one environment certifies only that environment; a gate born red and earned green is worth more than one that has never failed. Re-tuned, on evidence: the mechanical writing gates carried generic thresholds; the audit measured our actual published corpus — eight articles, twenty-four thousand words — and re-set them to the voice we actually publish, enforcing strictly what the corpus never does and tracking what it deliberately does. Some findings were deliberately banked rather than mass-fixed at speed, each with its reasoning recorded — the version of fast that still works a month later.
Not fixed, and named. The shared-blind-spot limit stands: a stale fact carried by both a producer and its challenger survived both — the challenger even flagged the correct value as the error; only an independent empirical check caught it. Paired review defends against producer error, not shared-context error. The single-arbiter dependency stands, and it is broader than one sentence usually admits: every contested finding resolves through one coordinating judgment, which is itself error-prone — three of its slips in the week of this article's production are logged in its production record — and which draws on the same model lineage as the agents it rules over. Four roles, one model family, plus one human. No external party has audited this audit; the checks are intra-system, and we regard commissioning that external review as a maturity step this paper's own standard eventually demands. The recurrence record stands: most failure classes were shifted or renamed rather than eliminated, one recurring inside the very session that named it. A defect class dies when a mechanical check makes it impossible, and several have not yet received that check. And the six-layer programme itself sits on this side of the ledger: its discharge is uneven — the health audit has run once but its cadence is not yet codified, the external hostile review is commissioned nowhere, and the codification layer's own specification is a tracked commitment rather than a document — now held on a layer-by-layer discharge ledger, so a ratified ambition cannot quietly become a memory. "Good enough" is not a state we recognise so much as a warning sign.
9. The Tools, Described
The platform hosts nine working applications behind shared two-factor authentication, and the audit recorded each one's maturity label as it stands — production-class where earned, prototype-class where true, including the estate's largest module: a record-keeping app for the practitioner's own working farm, the proving ground where new build disciplines meet real operating conditions before any client-facing use. The ninth of the nine is Fractal itself — deliberately the smallest code in the estate — the capability that answers the question every sole practitioner eventually faces: what happens if the machine dies? It carries a live three-copy backup model, a drilled recovery runbook, and one capability stated at exactly its verified level: designed and documented, not yet executed on real alternate hardware. Fractal's verification record includes its own worst day: the twin backup failures in Section 6 were failures of its earlier self, and the repair standard they produced — no backup is called restored until a restore drill has recovered a real artefact and verified the recovery against the source's own fingerprint — is now its standing acceptance test, with both backup tiers checked at every session open and a weekly sentinel sweeping the estate independently. It is the capability we would soonest be audited on, precisely because its record includes the day it failed and the controls that day bought.
The full paper carries the estate application by application, each with its audit-recorded maturity label, and the resilience capability set at its verified levels. Request the complete paper for the full account.
10. If You Are Buying Professional Work
The practical question this paper exists to sharpen: when any provider can produce polished, confident, fluent deliverables at negligible cost, what should a buyer of professional services actually ask for?
Our answer, from the inside: ask for the verification record. A tool list tells you what a practice bought. A methodology page tells you what it aspires to. A verification record tells you what its process catches — and a genuine one has recognisable features, because they are expensive to fake and uncomfortable to publish. It has numbers that are not flattering, like a zero-for-ten first-pass rate. It has catches in more than one direction, including against the practice's own coordinating judgment. It has named limits that have not been fixed. And it has dated, visible corrections on the public record, because a practice that corrects silently will eventually not correct at all. The 2025 case in the public record is instructive precisely because, on that record's account, the content passed the reviews that read for polish; the review that was missing was the one that checks claims against sources — the least glamorous and most consequential layer in this whole paper. In the vocabulary this practice was commissioned under, what a buyer is testing for is authenticity behind the presentation layer — and the verification record is where it either exists or does not.
We apply the ask to ourselves. This paper, its colophon, and the dated revision notes on this site are the public face of our verification record; this paper's own adversarial review is summarised in the colophon with its actual findings, and the address there reaches the practitioner directly. A claim in this paper challenged is a claim we re-verify on the record.
A note to fellow practitioners. Many sole founders are building AI-augmented practices right now, working these disciplines out from first principles, incident by incident — the same road this practice walked. Nothing in this paper is beyond a practice of one: every control described here was built by one practitioner, mostly in the week something went wrong. The advice this record supports is short. Write the check down, make it mechanical, and let it show you its evidence — whatever your equivalents of our gates turn out to be. The rest is refusing to let "it ran fine" stand in for "it was checked."
11. Where This Goes
An audit run once proves a practice could look. A standing cadence proves it keeps looking. The named remainder from Section 5 — the specification bodies, the decision ledgers, the operational archive, the engagement files under their confidentiality-first protocol — is the next review's scope, already written down, so that the follow-up starts from a list rather than a memory. The falsifier rule applies to the audit programme itself: if the next review finds nothing at all, the working assumption will be that the review has stopped seeing, not that the estate has stopped drifting.
This article closes its own loop. It was produced under the same governed method it describes — argument defined and approved before drafting; sources inventoried and gap-checked; external claims verified against their public record; the draft put through the same four-pass adversarial review, the same substance-traceability scoring with an independent human spot-check, and the same mechanical style gates it discusses — re-tuned, as Section 8 records, by the audit it reports. Its figures passed a pre-publication currency check that promptly caught a stale version reference in its own brief. The findings registers and evidence bases behind every number here persist as this paper's permanent companion record, which is what will make the next audit's first question answerable: what changed since we looked?
The premise we started with was that trust has to be designed in, because it cannot be inspected in at the end. The audit's answer is that designed-in trust is not a state a practice reaches. It is a discipline a practice runs — one that catches most things, misses some, names what it misses, and treats every miss as the specification for the next control. That is the record. We expect the next one to be better, and we have written down exactly what better would mean.
12. Bibliography
Associated Press (2025). Deloitte to partially refund Australian government for report with apparent AI-generated errors. AP wire, 7 October 2025.
CFO Dive — Alexis, A. (2025). Deloitte refunds over $60K for report with AI errors, Australian government says. CFO Dive, 21 October 2025.
Fishman, C. (1996). They Write the Right Stuff. Fast Company, Issue 06, December 1996/January 1997, p. 95.
Shao, Y., Jiang, Y., Kanell, T. A., Xu, P., Khattab, O., and Lam, M. S. (2024). Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models. Proceedings of NAACL 2024. Stanford Open Virtual Assistant Lab.
13. Colophon
About this article
Edition: First Edition — August 2026
How this article was produced
This article was produced under Steko's structured research-publication method, a governed production method for research articles. The method requires argument definition, source inventory with gap analysis, external citation verification, structured section architecture, adversarial review, and substance traceability verification before publication.
What the practitioner brought: Direction of the audit programme and this article's argument; the framing directives that shaped its register — humility carried by evidence rather than statement, findings brought forward for the time-limited reader, and figures detailed enough for an experienced peer to validate; independent verification of the three highest-stakes citations, the preparation for which surfaced a unit imprecision in one headline figure — a count reported as defects that was in fact a count of defective deliverables — before publication; the standing correction that verification steps are built into this practice, never optional extras; and every approval from story selection through publication.
What the production engine brought: The four audit streams and their paired adversarial verification; assembly of the evidence bases behind every section; draft production; verification of the external citations against their public record; a five-archetype hostile-reader review, substance-traceability scoring, register compliance review, mechanical style diagnostics, and a source-faithful resonance check — with the arbitration records of all of it persisted as this article's companion record.
Powered by Claude Fable 5
| Hostile reader review | 5 archetypes tested. The first round returned "needs revision" from all five — 21 objections, two assessed vulnerable. Every objection was resolved in the text or disclosed within it before publication. A confirmation round then verified the fixes and caught one further register defect in the corrections themselves, also fixed. |
| Substance traceability | First pass scored 80.7% — 46 of 57 factual claims verified, 4 found misstated against their sources and corrected, 7 traced to sources at arbitration. After correction, all 57 claims verified or corrected. One headline figure's unit was additionally tightened during spot-check preparation. |
| Practitioner spot-check | 3 citations independently followed to source by the practitioner. All confirmed. |
| Register compliance | Passed, with findings: one self-grading sentence and two hedging phrases identified and removed. The review's specific brief included catching performed humility as well as self-congratulation. |
| AI fingerprint mitigation | All enforced diagnostics passed at zero findings. Watch diagnostics measured and reported: em-dash density ~19 per 1,000 words, within the observed register of our published corpus; longest same-opener run 4; meta-discourse lexicon 0 hits. |
| Source-faithful resonance | 11 sections scored against their sources. One section took two revision rounds to move from "adequate" to "faithful" — its final wording is stricter, and better sourced, than the first draft's. |
| Practitioner parse + coverage round | After assembly, the practitioner's structured parse found three presentation defects no machine check had encoded and a register defect in the prose — and pressed the coverage question none of the gates had asked, forcing the mechanical check that proved a coverage gap against the commissioning brief: whole ruled elements, including both founding terms, had not reached the text. Publication was held, the article was revised at technical depth, and the miss bought standing controls: a requirements-traceability matrix, a dedicated coverage reviewer, and a rule that every build contract names its governing standards. This edition is the revised result, re-gated in full. |
We have made best efforts to ensure the accuracy and integrity of this article. Every external claim in it was verified against its cited public source before publication, and the highest-stakes citations were independently followed to source by the practitioner. We welcome refutation of anything published here — from any reader, and particularly from any party named in or affected by our citations. Write to thinking@steko.co.nz: a challenge that stands results in a visible, dated correction on this page, which is the practice this article describes.
The Audit We Ran on Ourselves · First Edition · August 2026 · steko.co.nz/thinking
The full paper carries the estate application by application, each with its audit-recorded maturity label, and the resilience capability set at its verified levels. Request the complete paper for the full account.
Request the full paper →