Agents not performing?
Can't run for hours without drifting?
Alignment.
FORGE WITH RIGOR · Jake Cukjati
Agent idle 30 min?
AGENT
0:00
You're the bottleneck.
Still prompting your agent to code?
87/87 agents, 12h37m+109,876 -19,451
112/112 agents, 16h28m+62,382 -3,845
Or handing off work for hours, days?
Prompt → code.
→
Specs → work.
Canonical.
Forge CTL
LLM inside a scaffold — Jack Clark
Forge CTL
One task at a time.
Gate every step.
Specify
Plan
Implement
The Bitter Lesson.
Build it. It gets replaced.
NEW
Claude Code · dynamic workflows
Forge CTL
Obsolete
Forge CTL → dynamic workflows (Agent SDK, pi-headless).
Model T assembly line
Not just an assembly line.
01
02
03
SHIP
01
02
03
SHIP
01
02
03
SHIP
01
02
03
SHIP
A coding factory.
Digital labor coordination.
01Specs
02Implementation
03Integration testing
04Code review
05Triage
Compound engineering
Spectacular
Specs by kind.
Store
Contract
Job
Adapter
Rule
Capability
Owed
Claimed
Closed
ROW
other
Nothing certifies its own work.
Forge
One agent at a time.
Gate every stage.
Plan
Implement
QA
E2E
Review
Archive
cap→ handoff →fresh
next but one
Rigor
Tests need a harness.
Harness = the asset.
harness
buildersstubsscenarios
Test fails? Cite the spec clause.
FAIL
Harness
spec clause
Stub
spec clause
App
spec clause
Test
spec clause
attempt
1 / 5
Reverted
Tests running.
No idea what's tested?
Who's in control?
Review the whole change.
diff
−
+
+
−
+
spec clause
Intent
Code quality
Code architecture
Infrastructure architecture
Alignment of objectives
Questions: why this way? How does it work?
Findings in 0 of 18.
53 findings 56 fixes ~13% of tokens
Compound.
Review → implementation references.
Two live sites, or it's a note.
DocumentWhat it coversWhen to use it
repository-constructionhow repos are builtnew data access
bulk-writesbatching, chunkingloading many rows
connection-poolingpool sizing, reusenew service, new client
candidate-a1 sitenote
candidate-b1 sitespromoted ✓
Plan cites itImplement obeys itReview grades against it
Repeat good patterns.
Triage.
Spec to-dos → what's next.
JEFF proposes. You apply.
§ 3.2
§ 1.4
§ 5.1
§ 2.7

Da Pile

Da WAAAGH!

Kunnin' Planz
ApplyDiscard
PLAN
Next lap: more accurate. HARNESS Plan Specs Implement Test Review Compound canonical Krumping · triage Spectacular Forge Rigor Forge review References
Every lap leaves context.
Then the prompts get tuned.
Open questions
Misaligned specs
Calibration log
References
jake-prompts
0 / 534
searched the whole disk for their own plugin.
one absolute path
PlanImplementReview… forge.js · cited Understand.
Harness-fy
Skill → CLI
SKILL.md
1
2
3
4
5
cli
draft
build
check
gate
$ harness next
state: review
findings: 0 open
STOP: USER_GATE
Always create CLIs for your skills.
Agents check specs.
Not each other.
Prompting AI to write code? Reconsider.
SPEC
Spec = canonical.
CLI your skills.
Compound every lap.
Open source — soon.
Thank you.
Jake Cukjati
307 contributions on September 25th
GitHub QR code
GitHub
X QR code
X