HAIL

How the model works

Learn is the seven heroes. About is the machinery behind them.

The idea

AI amplifies the system it enters — it doesn't fix a broken one. HAIL places a team on two axes and names the pattern so you can act.

Axes

X — AI Adoption Intensity: how much AI/agents are in the workflow (% AI code & PRs, AI review, agent-resolved tickets, agent count). High X is not “good” by itself — it only pays off if health can absorb it.

Y — Engineering System Health: the amplification coefficient. It decides whether AI-driven velocity compounds into value or into damage. Built from stability, quality, and security.

Vectors

Placement uses a six-dimension feature vector. Axes X and Y come from metric families below; the other four dimensions sharpen fit so volume, speed, review discipline, and the output–outcome gap can move a team between nearby heroes.

AI adoption

Feeds X

How aggressively AI and agents sit in the workflow — not whether that intensity is healthy. Signals include share of AI-authored or assisted code, share of majority-AI PRs, distinct agents merging to production, share of PRs getting automated AI review, and tickets resolved end-to-end by an agent.

Stability

Feeds Y · weight 40%

Does what you ship stay up? Change-failure rate and time-to-restore (MTTR) are the core tells; code churn can join when present. This is the largest slice of system health because instability is what AI tends to amplify first.

Quality

Feeds Y · weight 35%

Do defects escape, recur, or get rewritten soon after merge? Production defect density per 100 PRs, defect reopen rate, and short-window code churn measure whether velocity is buying rework. Quality falling while PR volume rises is a classic HOTSHOT / OVERLOAD signature.

Security

Feeds Y · weight 25%

Is AI-assisted code introducing exposure? New high/critical findings per 100 PRs and security-gate failure rate in CI capture whether completions and merges are adding OWASP-class risk faster than gates can catch it.

Productivity (volume)

Context vector · VOL

Raw output — merged PRs per developer per week and quarter-over-quarter PR-volume change. High volume alone never wins a placement; it only matters relative to health and outcomes.

Velocity

Context vector · VEL

Speed to delivered work: cycle time (first commit → merge), lead time (commit → deploy), and ticket-throughput change. Fast local cycles with flat ticket throughput is activity, not delivery.

Review health

Context vector · REVCOV

Is review real or rubber-stamp? Meaningful human-review coverage (and related unreviewed-merge pressure) is weighted above raw volume so the score resists gaming by flooding the pipeline.

Output–outcome gap

Context vector · GAP

The tell the model exists to expose: PR volume rising while ticket throughput does not keep pace. When activity climbs and value stays flat, GAP falls — the pattern behind OVERLOAD and other “busy but not shipping” diagnoses.

Honesty note: high output is not the goal. The gap between individual output and system outcome is the signal we care about most.

Placement
  1. Normalize — every metric 0–100 (100 = good end) using research anchors; high-is-bad metrics like failure rate are inverted.
  2. Build axes — X from adoption; Y = stability 40% + quality 35% + security 25%.
  3. Fit — weighted distance to seven archetype targets → fit % for all seven (health carries the most weight; review coverage above raw volume).
  4. Name — top archetype + transition path + actions. Confidence reflects how many metric families had real data.
Sources

Anchors and archetype hooks are paraphrased from published research — correlational and directional, not causal proof. Stats the model leans on:

  • DORA Report 2024 Stability tends to fall as AI use rises (−7.2% stability per +25% AI adoption in the reported relationship). About 39.2% of respondents report little or no trust in AI-generated code. Caution past strong foundations shows up as opportunity cost; low-health systems see throughput/stability tension.
  • DORA Report 2025 Core claim of the map: AI amplifies what is already there. Cluster profiles inform hero shapes — Harmonious High-Achievers (~20%), High Impact / Low Cadence (~7%), Legacy Bottleneck (~11%), Foundational Challenges (~10%), and Stable & Methodical. Elite engineering anchors include change-failure rate <5% and MTTR <1 hour (very weak restore times stretch toward multi-day).
  • METR randomized controlled trial Experienced developers were about 19% slower with AI tools while estimating they were about 20% faster — the HOTSHOT “feel fast / measure slow” pattern.
  • Faros AI engineering study Across 10,000+ developers and 1,255 teams: about +98% PRs merged, +91% review time, +154% PR size, with no company-level DORA improvement. Separately: ~31.3% more PRs merged without human review and roughly +9% per-developer bug rates — the OVERLOAD activity-without-outcome signature.
  • Veracode 2025 (GenAI code security) About 45% of AI completions introduce an OWASP Top 10 vulnerability; AI-assisted code showed about 2.74× more vulnerabilities than human-written code in the reported comparison. Feeds the security vector and HOTSHOT security hook.
  • GitClear Large-scale code-change analyses used for quality and churn framing — how often newly merged code is rewritten, and whether AI assistance correlates with more churn / copy-paste rather than durable change.
  • Stack Overflow Developer Survey Adoption, trust, and tooling-sentiment context for how teams actually use AI assistants — complements DORA’s trust and caution findings when placing low-adoption / high-health teams (STRONGHOLD) versus selective adopters (SHADOWWATCH).

Benchmarks on result cards attribute the specific study beside each comparison. Reassess quarterly; treat placement as a retrospective lens, not a ranking.

Team-level diagnostic for retrospectives — not an individual performance-ranking tool. Reassess quarterly.

Start now ▸
FAQ
What is HAIL?

HAIL (Heroes of the AI League) is a team diagnostic that scores AI adoption and engineering system health, then names the team's hero archetype with prioritized actions and a runbook.

What are the seven hero archetypes?

ALLOY MAN, SHADOWWATCH, STRONGHOLD, HOTSHOT, STORMLOCK, OVERLOAD, and KICK — each a distinct pattern of AI adoption versus engineering health, with its own primary risk and transition path. See Meet the Heroes.

Is HAIL free?

Yes — one scored diagnosis per browser is free, no email required (8 core questions, optional extend to 20; hero card, five actions, Save / Share / Download). The detailed What to Fix playbook and extra scored runs require Premium or Portfolio. Premium ($4.99) adds 3 extras over your existing result + advanced measurement; Portfolio ($49.99) adds 15 shared runs across named teams plus management; Enterprise is custom.

What does HAIL Premium include?

Premium costs $4.99. You get:

  • Three extra scored runs over your existing free result (Simple, Business/Professional, or Advanced)
  • Advanced real-measurement assessment (enter telemetry instead of estimation bands)
  • Unlock of the detailed What to Fix playbook on every result
  • Returning-user dashboard with run history
  • Activation email after purchase with a direct path to assess

One Premium per email. While extras remain, buy offers stay hidden; after all three are used, only Portfolio is offered.

Buy Premium ▸

What does HAIL Portfolio include?

Portfolio costs $49.99. Built for directors running more than one squad:

  • Shared pool of 15 scored runs across named teams (name + product/platform type)
  • Run / rerun in Simple, Professional, or Advanced (each burn one pool credit)
  • Management tabs: Teams, Compare, Org map, Insights, Members
  • Org map of latest Heroes + org-mean benchmarks
  • Insights rollup — averages, X/Y distribution, hero mix, movement, risk watchlist
  • History + sparklines, compare 2–3 teams, export / share
  • Invite emails that share the same pool; What to Fix unlocked

Use Portfolio in the top menu (or open portfolio.html) to buy or manage a shared pool.

Buy Portfolio ▸

Where do I see past runs?

Free results stay in this browser for 30 days — use Save for a private link (optionally emailed). After Premium or Portfolio, enter your email on the Assess page to retrieve your tests or Portfolio, or open the dashboard / Portfolio with your email.

Can AI agents run a HAIL assessment directly?

Yes — HAIL exposes list_questions, run_assessment, get_result, and get_playbook via MCP at heroleague.ai/mcp and via in-page WebMCP. Free run_assessment needs no email (rate-limited per IP). get_playbook requires Premium or Portfolio and a Save token proving ownership.