ENGINEERING HANDBOOK / FORGE LAB

Learn the system on the left. Grow one real project on the right.

Eighteen modules grow HeatStack Forge from a precise product contract into a recoverable, observable production agent. Software delivery is the required lab; office research and music creation prove the architecture can transfer.

No seven-day mastery. Build one piece of evidence today, then let the same system become stronger tomorrow.
RequiredSoftware deliveryTransferOffice researchTransferMusic creation

Locate the gap before choosing a start

Check only capabilities you can independently implement, explain, and verify. Results stay in this browser and travel with your exported local record.

First capability gap01 · Diagnostic & Evidence Map
Open chapter
HEATSTACK FORGE · 6 LAYERS · 18 MODULES

Six-layer agent system stack

Architecture answers how the parts work together; learning order answers what to study next. Both views use the same modules, URLs, and local progress.

Started
0/18
Tests run
0/18
Evidence verified
0/18
Portfolio-ready
0/18
Portfolio-ready currently means the module has locally validated lab evidence.
01

Foundations

Establish model inputs, context, and engineering foundations so the rest of the system starts from explicit, verifiable contracts.

This layer answers
What information does the model receive, and how does the developer create reliable inputs and engineering foundations?
Core concepts
AgentContextVerification & EvaluationObservability
MODULE01

Diagnostic & Evidence Map

Start from three real scenarios and define Forge users, boundaries, failing inputs, and its first acceptance evidence.

4 hours
Not started
Open chapter
MODULE02

AI Tools, Models & Task Judgment

Compare models and tools with benchmark tasks across quality, cost, latency, and context limits.

6 hours
Not started
Open chapter
MODULE03

Python, Git, HTTP & Engineering Foundations

Build the Forge CLI, configuration, schemas, HTTP client, async flow, logging, and tests.

10 hours
Not started
Open chapter
MODULE04

Model Mechanics & Multimodal Input

Understand tokens, context, embeddings, and vision input while building a unified input adapter.

10 hours
Not started
Open chapter
MODULE05

Prompts, Context & Structured Output

Use instruction hierarchy, context selection, few-shot examples, and schemas to build a stable requirement extractor.

10 hours
Not started
Open chapter
02

Core Agent Loop

Select actions from goals and state, execute tools, observe and verify results, then stop or replan.

This layer answers
How does the system choose actions, execute tools, observe results, and decide whether to stop or replan?
Core concepts
AgentToolVerification & Evaluation
MODULE07

Single-Agent & Tool-Calling Loop

Implement tool selection, structured arguments, streaming, retries, idempotency, and stop conditions.

12 hours
Not started
Open chapter
03

Capability Extensions

Extend what an agent can do through capability packages, external evidence, persistent state, and standard protocols.

This layer answers
How do we add Skills, retrieval, memory, and external protocol capabilities to an agent?
Core concepts
SkillRAGMemoryMCPContextTool
MODULE06

Skills, Prompts & Agent Packages

Package the requirement extractor as a Skill with a manifest, resource boundaries, and compatibility metadata.

9 hours
Not started
Open chapter
MODULE09

RAG, Retrieval & Knowledge Evaluation

Move from corpus ingestion, chunking, retrieval, reranking, and citations to an evaluated knowledge system.

14 hours
Not started
Open chapter
MODULE10

State, Sessions, Memory & Context Compression

Design session state, checkpoints, short- and long-term memory, compression, and replay.

12 hours
Not started
Open chapter
MODULE13

MCP Servers, SaaS & Streaming Protocols

Implement tools, resources, state, Streamable HTTP, authentication, and client integration.

16 hours
Not started
Open chapter
04

Workflows & Orchestration

Decompose complex work into recoverable steps and coordinate execution units through explicit contracts.

This layer answers
How do we decompose complex tasks and coordinate steps or multiple agents?
Core concepts
PlanningSub-agentAgentVerification & Evaluation
MODULE11

Planner/Executor & Agentic Workflows

Build task graphs, planning, execution, validation, replanning, budgets, and loop protection.

14 hours
Not started
Open chapter
MODULE12

Multi-Agent Collaboration & Orchestration

Use message contracts, shared state, conflict handling, and cost evaluation to decide when roles should split.

14 hours
Not started
Open chapter
05

Harness & Governance

Use the runtime environment to control tools, permissions, side effects, budgets, human approval, recovery, and platform differences.

This layer answers
What controls permissions, runtime boundaries, side effects, platform adaptation, and recovery?
Core concepts
HarnessToolSkillObservabilityVerification & Evaluation
MODULE08

Tools, Sandboxing & Side-Effect Control

Build tool registration, dry runs, change plans, confirmation, rollback, and audit logs.

12 hours
Not started
Open chapter
MODULE14

Identity, Permissions, Security & Supply Chain

Cover identity, authorization, least privilege, human confirmation, injection attacks, dependency risk, and incident response.

14 hours
Not started
Open chapter
MODULE15

Cross-Platform Adapters & Migration

Build capability, directory, and permission adapters for Codex, Claude Code, and WorkBuddy.

12 hours
Not started
Open chapter
06

Production Engineering

Use evaluation, observability, deployment, and evidence to prove the system is reliable under real constraints.

This layer answers
How do we prove reliability and complete deployment, portfolio evidence, and interview communication?
Core concepts
Verification & EvaluationObservabilityHarness
MODULE16

Evaluation, Observability, Deployment & Reliability

Build eval sets, tracing, quality and cost metrics, deployment, rollback, and incident drills.

18 hours
Not started
Open chapter
MODULE17

End-to-End Portfolio Delivery

Integrate all three scenarios and complete the release, demo, architecture docs, and evidence index.

24 hours
Not started
Open chapter
MODULE18

Interview, System Design & Industry Judgment

Use Forge code, logs, evals, and failure evidence to practice project explanation and system design.

10 hours
Not started
Open chapter

Practice queue0

Turn selected Skills into precise course stages, lab evidence, and interview claims.

The queue is empty. Choose one capability you genuinely want to practice from a Skill detail page.

LOCAL PORTFOLIO / BROWSER ONLY

Learning record and portfolio evidence

Export diagnostic answers, the practice queue, progress, and lab evidence. Import and validation happen only in this browser; the file is never uploaded. Portfolio-ready means locally validated lab evidence is available.

Started
0
Tests run
0
Evidence verified
0
Portfolio-ready
0