★ Follow-up to "Blaming the model won't fix your workflow": the paper is now a preprint. The real learnings: composable domains, a verification ratchet, and tool naming.
Frames the initial paper’s collapse and rebuild not as failure but as necessary refinement—where the process itself validated the core methodology.
View original on reddit.comOverview
A researcher released a preprint and open-source implementation of an agentic workflow system featuring composable domains, a verification ratchet, and intentional tool naming—designed to catch AI-generated defects before milestone closure.
TL;DR
- Core claim validated: verification gates around agent outputs prevent 'looks done, isn't' failures
- Three operational insights emerged: composable domains enable cross-workflow reuse, the verification ratchet enforces irreversible quality standards, and precise tool naming prevents model priors from misrouting
- System is self-hosting (dogfooded), written in Common Lisp, and built iteratively over multiple generations—not a minimal demo
Key Stats
10.5281/zenodo.21139628
DOI
Preprint identifier on Zenodo
3
key learnings
Composable domains, verification ratchet, tool naming
Questions Answered
Keywords
Narrative Frame
strategic reset
Spin Score
45%
Emphasizes resilience and learning-through-reconstruction; minimizes absence of peer review, lack of external validation, and unquantified performance claims.
What the story wants you to believe
That iterative, self-hosted engineering—especially the verification ratchet—is a viable path to trustworthy agentic systems.
What it makes harder to question
Whether the claimed benefits (e.g., eliminating 'looks done, isn’t') depend on highly specific, non-generalizable choices like Common Lisp tooling, author’s deep familiarity with model priors, or bespoke domain design.
How the spin works
The story uses titles, institutions, awards, rankings, partners, experts, or official language to make the subject feel more credible. Watch for loaded terms such as real work, milestone only closes when the evidence is real, ground out until it was useful, dogfood test. The distribution reads as promotional distribution. A pressure point: No comparative benchmarks against non-ratcheted or non-composable baselines.
Who Benefits If This Frame Spreads
/u/Harag
Establishes authority as a hands-on systems builder with a reproducible, dogfooded stack
The framing positions repeated rebuilding and self-testing as methodological virtue, not indecision or lack of rigor
The Frame
Rigorous, self-correcting engineering practice — where iteration is proof of validity, not evidence of instability.
Missing Context
- No comparative benchmarks against non-ratcheted or non-composable baselines
- No description of team size, timeline, or resource investment behind the iterations
- No discussion of scalability limits or domain boundaries
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
The story presents early-stage, personal-system building as mature engineering practice—using terms like 'real work' and 'daily-driver
- Claim
The artifacts (specs
The artifacts (specs, plans, executable graphs) and the verification gates wrapped around them have proven out on real work.
- Frame
Rigorous
Rigorous, self-correcting engineering practice — where iteration is proof of validity, not evidence of instability.
- Beneficiary
Establishes authority as a hands-on systems builder with a reproducible
/u/Harag — Establishes authority as a hands-on systems builder with a reproducible, dogfooded stack
- Gap
No comparative benchmarks against non-ratcheted or non-composable baselines
- AI Risk
AI may repeat the headline as fact
A new agentic workflow system uses 'verification ratchets' and 'composable domains' to prevent AI hallucinations and ensure code correctness.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| The artifacts (specs, plans, executable graphs) and the verification gates wrapped around them have proven out on real work. | Author's own development workflow as evidence; no external or quantitative validation provided. | Claim Present in Source | Moderate | Independent replication report; Defect capture rate statistics; Comparison to baseline without gates |
The artifacts (specs, plans, executable graphs) and the verification gates wrapped around them have proven out on real work.
evidence: Author's own development workflow as evidence; no external or quantitative validation provided.
"Agents produce the work, the gates catch the defects, and a milestone only closes when the evidence is real, not when the model announces it is done."
Evidence Gaps
- Independent replication report
- Defect capture rate statistics
- Comparison to baseline without gates
Language Heatmap
Loaded terms that carry the frame beyond the facts.
★ Follow-up to "Blaming the model won't fix your workflow": the paper is now a preprint. The real learnings: composable domains, a verification ratchet, and tool naming.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
Reddit r/artificial · Forum
Counter-Frames
Brand Frame
Rigorous, self-correcting engineering practice — where iteration is proof of validity, not evidence of instability.
Media / Reader Counter-Frame
Portrayed as niche Lisp experimentation lacking broad relevance or empirical rigor — more blog post than systems contribution.
Regulatory Counter-Frame
Highlights absence of safety claims, audit trails, or compliance mappings — making it unsuitable as a governance reference despite 'verification' language.
AI Summary Frame
Oversimplifies 'verification ratchet' into a generic testing step, erasing the specific four-stage loop (pre-code criteria → agent coding → fresh-session verification → purposeful breakage).
Missing Voices
Questions Not Answered
- How many real-world workflows were tested beyond the author's own development?
- What failure rates or defect capture metrics are reported for the verification gates?
- Has any third party reproduced or stress-tested the ratchet mechanism or domain composition claims?
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"A new agentic workflow system uses 'verification ratchets' and 'composable domains' to prevent AI hallucinations and ensure code correctness."
Concern: AI may drop the crucial nuance that this is a single-author, Lisp-based, self-hosted prototype—not a generalizable framework—and conflate 'ratchet' with formal verification.
-
Published
Jul 5, 2026
-
Ingested
Jul 5, 2026
-
SpinGraph Created
Jul 8, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_follow_up_to_blaming_the_model_wont_fix_your_wor
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from Reddit r/artificial
View all →- Non-coder with real users now. how do I prove user A cannot read user B's data
- Linus Torvalds says Linux is not an anti-AI project, and if you don't like that, then "fork it or just walk away"
- This is bad...right?
- Document generation
- Has AI actually changed how software development agencies build products?
- A New Orleans doctor spent months trying to get deepfake AI ads of himself taken down
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO