Agentic Labs Build with us
Agentic Labs field guide / 002

The agent
engineering manual.

Six tactical patterns that turn internal AI from a clever local experiment into shared, observable, self-improving infrastructure.

By Max KnoxAugust 18, 202616 min read
agent-runtime / trace_1842live

09:41:02 objective.received

"Which founder cohorts improved after office hours?"

09:41:03 schema.read models/cohorts.sql

09:41:03 tool.call readonly.query

SELECT cohort, delta, evidence
FROM agent_context.cohort_outcomes
WHERE event_type = 'office_hours';

09:41:04 retrieval.fuse sql + graph + semantic

09:41:05 evidence.rerank 12 → 5 passages

09:41:06 response.complete grounded

5 sources cited842 msread-only
00 / The engineering thesis

Agent capability grows when you give models better primitives, not more predetermined paths.

The practical challenge is not making an agent perform one impressive task. It is building an environment where agents can discover context, choose reliable tools, preserve successful workflows, learn from traces, and serve an entire organization without becoming an ungovernable pile of prompts.

01 / Pattern index
Pattern 01 / Context extraction

Give the agent a map
and a read-only key.

Two low-level tools cover a surprising amount of analytical work: inspect the current schema, then query the data directly.

A common instinct is to pre-build a narrow function for every question the business might ask. That creates a queue: every new analysis needs engineering to invent another endpoint. A capable agent can instead inspect model definitions, understand relationships, and construct the query required for the question in front of it.

Direct does not mean unrestricted. The database role should be incapable of mutation, limited to approved schemas or replicas, protected by statement and row limits, and fully logged. The model gets broad analytical flexibility inside a hard deterministic boundary.

Context tools
Active roleagent_analystSELECT only
models/revenue/cohort_outcomes.sqlschema context
-- active model definition
MODEL cohort_outcomes {
  primary_key: cohort_id
  dimensions: [program, batch, region]
  measures: [revenue_delta, retention_90d]
  joins: {
    office_hours: cohort_id -> cohort_id
    interactions: founder_id -> founder_id
  }
}
WHY IT MATTERS

The agent sees the same active definitions the application uses, reducing invented joins and stale assumptions.

01Replica or approved views

Keep operational writes physically or logically out of reach.

02Database-enforced role

Do not rely on a prompt to prohibit mutation.

03Budgets and timeouts

Limit rows, runtime, concurrency, and export size.

04Trace every query

Store identity, purpose, SQL, result size, and latency.

Pattern 02 / Agent context store

Normalize for truth.
Denormalize for retrieval.

The operational schema protects integrity. The agent-facing schema should reduce the number of joins required to understand an entity or event.

Fragmented sources
CRMMeetingsBillingSupport
→
Normalizecanonical IDs
typed events
→
Denormalizeentity context
one retrieval row
→
Indexlexical + vector
graph edges
→
Retrievefuse + rerank
cite evidence
Retrieval stack
RRF

Reciprocal Rank Fusion combines ranked lists without requiring their raw scores to be comparable.

score(d) = Σ 1 / (k + ranki(d))
QUERY

Why did the robotics cohort improve after June?

01 / meeting transcriptOffice-hours format changed to live teardown sessions
relevance 96%
semantic + keyword + graph
02 / cohort metricsActivation rose 18% within two weeks of the format change
relevance 91%
keyword + graph
03 / founder interactionsRepeated deployment blockers declined across six teams
relevance 84%
semantic + graph

Four signals active. Results are fused, reranked, and ready for citation.

Pattern 03 / Skill lifecycle

When a workflow works,
compile the lesson.

A successful ad-hoc session should be able to graduate into a reusable organizational capability without becoming a copy-pasted prompt.

session / investigation_214AD-HOC RUN
EXECUTE FIRST

Solve a concrete task in the interactive harness.

The operator and agent investigate a real problem together. The trace captures tools, context, corrections, decisions, and the final successful output.

  • Keep the complete execution trace
  • Mark human corrections and judgment calls
  • Confirm the outcome is actually useful
$agent run investigation_214 --trace
Design principleSkills are not chat transcripts with a filename. A durable skill separates objective, inputs, procedure, policy, exceptions, and evaluation.
Pattern 04 / Resolver hygiene

Lint the capability layer
like production code.

As skills accumulate, routing quality becomes an engineering concern. Overlap creates indecision; gaps create improvisation.

check_resolvable / agents.md
skillreschedule_office_hourscalendar changes for group sessions
skillmove_founder_sessionmove an office-hours calendar event
skillanalyze_cohort_outcomescompare cohort performance by intervention
routecustomer_communication.*send, cancel, or revise a customer message
>_

Run the resolver check to inspect overlap, coverage, names, and parameter boundaries.

DRYOverlap → parameterize

One canonical skill handles the shared procedure with explicit variants.

+
MECEAmbiguity → separate

Each route owns a distinct intent; together they cover the domain.

=
RESOLVABLEOne clear entry point

The planner can select a capability without semantic coin flips.

Pattern 05 / Autonomous evaluation

Let the system study
its own working day.

A nightly evaluator can turn traces and meeting artifacts into proposed improvements, while version control and human review keep production changes deliberate.

18:0022:0002:0006:00
INGESTToday's artifacts
  • 184 agent interactions
  • 7 expert meetings
  • 23 human corrections
  • 11 failed tool calls
ANALYZEPattern extraction
  • Cluster repeated failures
  • Compare expert outputs
  • Detect missing context
  • Generate regression cases
PROPOSEReviewable changes
  • 3 skill patches
  • 2 schema descriptions
  • 1 resolver merge
  • 8 CRM facts captured
02:14:08evaluation queue ready / 225 artifacts indexed
Critical boundary

Autonomous evaluation does not require autonomous deployment.

Let evaluators produce patches, evidence, and tests. Require review before changing production prompts, tools, permissions, or records that trigger downstream action. The learning loop can move quickly without becoming self-authorizing.

Pattern 06 / Shared execution

Graduate from a smart
terminal to team infrastructure.

Local harnesses are ideal for discovery. Shared runtimes are where context, skills, observability, and organizational learning compound.

LOCAL DISCOVERY

One operator. One harness. Fast learning.

A local CLI gives a technical operator maximum flexibility to discover useful workflows. The limitation is not capability; it is distribution.

  • Fastest path to experimentation
  • Private context and local skills
  • Knowledge leaves when the terminal closes
JUST-IN-TIME SOFTWARE

Build the answer, not another permanent feature.

Once the shared runtime has data and tools, a user can request the exact report or transient interface needed for a decision. Keep the generated artifact when it becomes recurring; discard it when it has served its purpose.

>

Submit the request to assemble a temporary interface.

07 / Implementation runbook

Ship the substrate
in four sprints.

Runbook progress0 / 4

Build one vertical slice through context, tools, skills, evaluation, and shared access. Avoid launching five disconnected experiments that cannot teach each other.

Sprint 01Context surface

Make one business domain legible.

Choose a high-value analytical domain. Publish current schema definitions, create the constrained read role, and prepare an entity-level context table for retrieval.

  • Schema reader tool
  • Read-only query tool
  • Canonical IDs and events
  • Query audit log
Acceptance test

The agent answers ten representative questions with executable SQL, cited evidence, and zero mutation capability.

Sprint 02Retrieval stack

Make the right evidence easy to find.

Index the denormalized context through lexical, semantic, and relationship-aware retrieval. Fuse ranked lists, rerank the short list, and preserve source provenance.

  • BM25 index
  • Vector index
  • Graph expansion
  • Retrieval evaluation set
Acceptance test

Known-answer queries retrieve the required source in the top five, and every generated claim can point back to evidence.

Sprint 03Skills + resolver

Turn successful sessions into maintained capabilities.

Define the skill template, package the first proven workflows, register clear routes, and run overlap and coverage checks before merge.

  • Versioned skill format
  • Skill creation meta-flow
  • Resolver registry
  • DRY / MECE checks
Acceptance test

A teammate can invoke the capability from a fresh session without knowing its tool sequence or original author.

Sprint 04Shared learning loop

Centralize access and install nightly evaluation.

Move validated capabilities into a shared harness, capture traces by default, and schedule an evaluator that proposes versioned improvements with regression tests.

  • Shared identity and access
  • Central trace store
  • Nightly evaluation job
  • Human review queue
Acceptance test

A correction from one teammate produces a tested improvement proposal that can benefit every user after review.

08 / Production checklist

The demo is done.
Is the system ready?

Use these questions before expanding access or autonomy.

DataCan the database enforce every promised access boundary?
RetrievalDo you measure whether required evidence is actually found?
SkillsCan procedures be versioned, tested, and rolled back?
RoutingDoes each intent have one obvious capability?
EvaluationAre suggested changes paired with evidence and regression tests?
RuntimeCan every action be traced to a user, tool, policy, and result?
09 / The compounding advantage

Build fewer features.
Build better primitives.

The strongest internal agent systems do not predict every workflow in advance. They create a controlled environment where models can inspect context, compose reliable tools, and preserve what works.

That is the shift from automation projects to organizational infrastructure: every successful investigation can become a skill, every correction can become an evaluation, and every teammate can inherit what the system learned yesterday.

Continue the series

The AI-Native Organization Playbook

Zoom out from the engineering patterns to the operating model, culture, and rollout strategy around them.

Read field guide 001 →
Engineer the first vertical slice

Agentic Labs designs internal agent systems around your real data, tools, and operating constraints.

Talk to Agentic Labs