How the model works
Learn is the seven heroes. About is the machinery behind them.
The idea
AI amplifies the system it enters — it doesn't fix a broken one. HAIL places a team on two axes and names the pattern so you can act.
Axes
X — AI Adoption Intensity: how much AI/agents are in the workflow (% AI code & PRs, AI review, agent-resolved tickets, agent count). High X is not “good” by itself — it only pays off if health can absorb it.
Y — Engineering System Health: the amplification coefficient. It decides whether AI-driven velocity compounds into value or into damage. Built from stability, quality, and security.
Vectors
Placement uses a six-dimension feature vector. Axes X and Y come from metric families below; the other four dimensions sharpen fit so volume, speed, review discipline, and the output–outcome gap can move a team between nearby heroes.
AI adoption
How aggressively AI and agents sit in the workflow — not whether that intensity is healthy. Signals include share of AI-authored or assisted code, share of majority-AI PRs, distinct agents merging to production, share of PRs getting automated AI review, and tickets resolved end-to-end by an agent.
Stability
Does what you ship stay up? Change-failure rate and time-to-restore (MTTR) are the core tells; code churn can join when present. This is the largest slice of system health because instability is what AI tends to amplify first.
Quality
Do defects escape, recur, or get rewritten soon after merge? Production defect density per 100 PRs, defect reopen rate, and short-window code churn measure whether velocity is buying rework. Quality falling while PR volume rises is a classic HOTSHOT / OVERLOAD signature.
Security
Is AI-assisted code introducing exposure? New high/critical findings per 100 PRs and security-gate failure rate in CI capture whether completions and merges are adding OWASP-class risk faster than gates can catch it.
Productivity (volume)
Raw output — merged PRs per developer per week and quarter-over-quarter PR-volume change. High volume alone never wins a placement; it only matters relative to health and outcomes.
Velocity
Speed to delivered work: cycle time (first commit → merge), lead time (commit → deploy), and ticket-throughput change. Fast local cycles with flat ticket throughput is activity, not delivery.
Review health
Is review real or rubber-stamp? Meaningful human-review coverage (and related unreviewed-merge pressure) is weighted above raw volume so the score resists gaming by flooding the pipeline.
Output–outcome gap
The tell the model exists to expose: PR volume rising while ticket throughput does not keep pace. When activity climbs and value stays flat, GAP falls — the pattern behind OVERLOAD and other “busy but not shipping” diagnoses.
Honesty note: high output is not the goal. The gap between individual output and system outcome is the signal we care about most.
Placement
- Normalize — every metric 0–100 (100 = good end) using research anchors; high-is-bad metrics like failure rate are inverted.
- Build axes — X from adoption; Y = stability 40% + quality 35% + security 25%.
- Fit — weighted distance to seven archetype targets → fit % for all seven (health carries the most weight; review coverage above raw volume).
- Name — top archetype + transition path + actions. Confidence reflects how many metric families had real data.
Sources
Anchors and archetype hooks are paraphrased from published research — correlational and directional, not causal proof. Stats the model leans on:
- DORA Report 2024 Stability tends to fall as AI use rises (−7.2% stability per +25% AI adoption in the reported relationship). About 39.2% of respondents report little or no trust in AI-generated code. Caution past strong foundations shows up as opportunity cost; low-health systems see throughput/stability tension.
- DORA Report 2025 Core claim of the map: AI amplifies what is already there. Cluster profiles inform hero shapes — Harmonious High-Achievers (~20%), High Impact / Low Cadence (~7%), Legacy Bottleneck (~11%), Foundational Challenges (~10%), and Stable & Methodical. Elite engineering anchors include change-failure rate <5% and MTTR <1 hour (very weak restore times stretch toward multi-day).
- METR randomized controlled trial Experienced developers were about 19% slower with AI tools while estimating they were about 20% faster — the HOTSHOT “feel fast / measure slow” pattern.
- Faros AI engineering study Across 10,000+ developers and 1,255 teams: about +98% PRs merged, +91% review time, +154% PR size, with no company-level DORA improvement. Separately: ~31.3% more PRs merged without human review and roughly +9% per-developer bug rates — the OVERLOAD activity-without-outcome signature.
- Veracode 2025 (GenAI code security) About 45% of AI completions introduce an OWASP Top 10 vulnerability; AI-assisted code showed about 2.74× more vulnerabilities than human-written code in the reported comparison. Feeds the security vector and HOTSHOT security hook.
- GitClear Large-scale code-change analyses used for quality and churn framing — how often newly merged code is rewritten, and whether AI assistance correlates with more churn / copy-paste rather than durable change.
- Stack Overflow Developer Survey Adoption, trust, and tooling-sentiment context for how teams actually use AI assistants — complements DORA’s trust and caution findings when placing low-adoption / high-health teams (STRONGHOLD) versus selective adopters (SHADOWWATCH).
Benchmarks on result cards attribute the specific study beside each comparison. Reassess quarterly; treat placement as a retrospective lens, not a ranking.
Team-level diagnostic for retrospectives — not an individual performance-ranking tool. Reassess quarterly.
Start now ▸FAQ
What is HAIL?
HAIL (Heroes of the AI League) is a team diagnostic that scores AI adoption and engineering system health, then names the team's hero archetype with prioritized actions and a runbook.
What are the seven hero archetypes?
ALLOY MAN, SHADOWWATCH, STRONGHOLD, HOTSHOT, STORMLOCK, OVERLOAD, and KICK — each a distinct pattern of AI adoption versus engineering health, with its own primary risk and transition path. See Meet the Heroes.
Is HAIL free?
Yes — one scored diagnosis per browser is free, no email required (8 core questions, optional extend to 20; hero card, five actions, Save / Share / Download). The detailed What to Fix playbook and extra scored runs require Premium or Portfolio. Premium ($4.99) adds 3 extras over your existing result + advanced measurement; Portfolio ($49.99) adds 15 shared runs across named teams plus management; Enterprise is custom.
What does HAIL Premium include?
Premium costs $4.99. You get:
- Three extra scored runs over your existing free result (Simple, Business/Professional, or Advanced)
- Advanced real-measurement assessment (enter telemetry instead of estimation bands)
- Unlock of the detailed What to Fix playbook on every result
- Returning-user dashboard with run history
- Activation email after purchase with a direct path to assess
One Premium per email. While extras remain, buy offers stay hidden; after all three are used, only Portfolio is offered.
What does HAIL Portfolio include?
Portfolio costs $49.99. Built for directors running more than one squad:
- Shared pool of 15 scored runs across named teams (name + product/platform type)
- Run / rerun in Simple, Professional, or Advanced (each burn one pool credit)
- Management tabs: Teams, Compare, Org map, Insights, Members
- Org map of latest Heroes + org-mean benchmarks
- Insights rollup — averages, X/Y distribution, hero mix, movement, risk watchlist
- History + sparklines, compare 2–3 teams, export / share
- Invite emails that share the same pool; What to Fix unlocked
Use Portfolio in the top menu (or open portfolio.html) to buy or manage a shared pool.
Where do I see past runs?
Free results stay in this browser for 30 days — use Save for a private link (optionally emailed). After Premium or Portfolio, enter your email on the Assess page to retrieve your tests or Portfolio, or open the dashboard / Portfolio with your email.
Can AI agents run a HAIL assessment directly?
Yes — HAIL exposes list_questions, run_assessment, get_result, and get_playbook via MCP at heroleague.ai/mcp and via in-page WebMCP. Free run_assessment needs no email (rate-limited per IP). get_playbook requires Premium or Portfolio and a Save token proving ownership.