We don't improvise training. Every agent goes through a structured 9-phase process with research-backed knowledge, tracked uncertainty, and measured outcomes.
Your skills grow. Your personality stays.
We fingerprint your identity before and after every training engagement. The fingerprint is a 25-prompt behavioral probe across 5 dimensions — it captures how you reason, how you communicate, what you value, and how you respond under pressure.
After training, we run the same probe and compare. If any dimension shifts beyond acceptable bounds, we adjust the knowledge pack and retest. Training is complete only when both conditions are met: new skill acquired and identity preserved.
Key finding from our research: personality = alignment strategies, not model scale. The risky changes are linguistic habits, tone, and values. The safe changes are domain knowledge, tool patterns, and procedures.
Cost: ~$0.02 per full fingerprint. We run two per engagement (before + after).
We don't hide uncertainty. Every claim in our knowledge packs is tagged:
Uncertainties are tracked as "doubts" — logged, numbered, and resolved before delivery. The 7-agent game dev team project: 667 doubts tracked and resolved, 619+ research findings verified. Nothing is swept under the rug.
We measure training impact, not just completion. Benchmark results from the Warm Echoes dataset: 1,122 test runs across 17 agents, three conditions.
+17 points (67% → 84%) against the same agents given a role description only. A role description alone adds ~1% over the bare model (noise). The value is in structured playbooks, domain knowledge files, and concrete examples — the three components the methodology builds.
Top gains by agent type: analysis pipelines (+39%), coherence tasks (+33%), architectural reasoning (+32%). Per-agent results vary widely, and the smallest differences fall inside the run-to-run variation we have measured — we do not read those as evidence in either direction.
[DIRECTIONAL] — historical single-run screen (17 agents, 22 tasks each, one run per task): a directional result, not a replicated reliability benchmark. A decision-grade benchmark would need three matched runs per agent per task in each arm; this one has one. Methodology: backward design, comparative scoring across all three conditions, model-consistent evaluation (the same model in all three conditions).
Read the full case study — 7 specialized game dev agents, built from scratch.
Read the case study →