The Curator

How to Build an AI Capability Radar Before the Market Makes the Shift Obvious

Last updated: 10/9/2026

Back to blog
Daniel Rosenthal avatarDaniel Rosenthal 7 min read
Cover image for How to Build an AI Capability Radar Before the Market Makes the Shift Obvious
AI-assisted, human-reviewed. Drafted with AI research tools from public sources and edited by our team. How we build these →

Most teams monitor artificial intelligence through announcements: a new model, benchmark, interface, or agent framework appears, and someone forwards the link. This creates awareness without direction. The collection grows, but no one can explain which change matters, what it makes possible, or when to act.

A capability radar is different. It tracks what systems can reliably do, under which conditions, and what product assumptions those capabilities may overturn. The concrete outcome is a living portfolio of evidence-backed opportunities, each connected to a small experiment and a decision date.

Step 1: Define the decision your radar must improve

Begin with a decision, not a technology category. “Track generative AI” is too broad. “Identify when support workflows can move from drafting to supervised resolution” is useful because it establishes a destination and a threshold.

Write one decision question using this form: When can we safely replace or redesign a specific step for a specific user under specific constraints?

For example: “When can an AI system reconcile routine supplier invoices while sending ambiguous cases to an operator?” This directs attention toward document variability, tool access, error detection, latency, permissions, and escalation. A leaderboard score alone cannot answer it.

Common mistake: choosing a topic instead of a decision

A topic produces an archive. A decision produces criteria. If every interesting release qualifies for the radar, the scope is still too loose.

Step 2: Decompose the workflow into capability claims

Map the present workflow as observable actions. Avoid broad labels such as “reasoning” or “automation.” For invoice reconciliation, the actions might be:

  1. Extract supplier, amount, currency, date, and purchase-order reference.
  2. Retrieve the matching order and delivery record.
  3. Compare values while applying tolerance rules.
  4. Identify missing or contradictory evidence.
  5. Post an approved transaction or assemble an exception packet.
  6. Record what happened and why.

Turn each action into a falsifiable capability claim. “The system can retrieve the correct purchase order from incomplete identifiers” is testable. “The system understands procurement” is not.

This decomposition reveals where innovation is actually needed. A stronger model may improve extraction, while the decisive obstacle could be identity resolution, authorization, or recovery after a tool failure.

Common mistake: treating the model as the whole system

Useful capabilities emerge from models, retrieval, interfaces, policies, tools, and human review. Track the weakest dependency rather than crediting every outcome to the model.

Step 3: Create a radar with stages that imply action

Use stages based on evidence and commitment. Names matter less than explicit entry rules.

StageEvidence requiredTeam action
SignalA credible demonstration or technical change suggests a new capability.Record the claim and identify what remains unproven.
ProbeThe capability can be tested with representative internal examples.Run a bounded evaluation without production access.
PilotThe system meets a defined threshold and failure containment is plausible.Use it with limited users, permissions, and reversible actions.
OperationalPerformance, economics, governance, and recovery work under real conditions.Integrate it into the workflow and monitor drift.
ParkedA named constraint blocks progress.Revisit only when evidence changes that constraint.

Movement between stages should require evidence. A persuasive demo can create a signal, but it cannot authorize a pilot. A successful laboratory test can justify a probe, but it says little about production permissions or exception handling.

Common mistake: using vague horizons

Labels such as “now,” “next,” and “later” conceal why an item sits where it does. Evidence-based stages make disagreement productive: participants can debate the missing proof rather than speculate about timing.

Step 4: Build an evidence card for every capability

Each radar item needs a compact record. Without one, memory gradually converts demonstrations into facts.

  • Capability claim: The exact behavior being assessed.
  • Workflow affected: The user and step that could change.
  • Evidence: What was observed, with reproducible inputs where possible.
  • Boundary conditions: Languages, formats, tools, permissions, and context required.
  • Failure modes: How the capability breaks and whether failure is detectable.
  • Dependency: The model, data, infrastructure, policy, or interface it relies on.
  • Next test: The cheapest experiment that could change the decision.
  • Review trigger: An event or date that prompts reassessment.

Suppose a system correctly reconciles clean invoices in a demonstration. The evidence card should not say “invoice automation works.” It might say: “On digitally generated invoices from known suppliers, the system extracted required fields and matched exact purchase-order references. Scanned documents, partial deliveries, duplicate invoices, and conflicting currencies remain untested.” That distinction protects the team from premature certainty.

Common mistake: recording conclusions without conditions

Capability is conditional. Always preserve the environment in which success occurred. Remove those conditions and the radar becomes marketing material.

Step 5: Design tests around consequential failure

A useful evaluation does not merely sample typical inputs. It represents the distribution the system will face and deliberately includes cases where an incorrect action would matter.

Create three sets. The routine set covers frequent, well-formed cases. The edge set includes rare formats, missing fields, conflicting records, and tool interruptions. The adversarial set tests misleading instructions, manipulated documents, unauthorized requests, or poisoned retrieved content.

Define success at the workflow level. For reconciliation, measure whether the final disposition is correct, whether the supporting evidence is complete, whether unsafe actions are blocked, and whether uncertain cases reach the right reviewer. A field-level extraction score can hide a financially consequential mismatch.

Use a shadow run before granting write access. Let the system process real or appropriately protected historical cases, but compare its proposed actions with completed outcomes. Then inspect disagreements. Some will be model failures; others will expose inconsistent human rules or poor source data.

Common mistake: averaging away dangerous errors

An acceptable overall result can conceal a severe failure class. Report results by consequence: harmless formatting errors, recoverable routing errors, unauthorized actions, and silent incorrect decisions should not share one undifferentiated measure.

Step 6: Translate technical movement into product options

Do not ask only, “Can this task now be automated?” A new capability can support several product moves:

  • Compress: reduce the time required for an existing step.
  • Reorder: perform analysis earlier, before a user commits effort.
  • Personalize: adapt the workflow to context that static software ignored.
  • Unbundle: separate expert judgment from routine preparation.
  • Eliminate: remove a coordination step rather than making it faster.

In the invoice example, extraction may merely compress data entry. Reliable exception classification could reorder work by showing reviewers the riskiest cases first. Evidence assembly might unbundle investigation from approval. The larger opportunity is often a redesigned workflow, not an artificial intelligence button attached to the old one.

Common mistake: mistaking feasibility for value

A technically possible feature may add review work, increase liability, or solve a low-friction problem. Pair every capability claim with a user consequence and an operational consequence.

Step 7: Run a disciplined review cadence

Assign one owner to each radar item and review only changes in evidence. Ask four questions: What became newly possible? What constraint disappeared? What failure appeared? What experiment should follow?

Move an item forward only when its entry rule is satisfied. Park it with a named blocker such as “cannot detect duplicate documents across subsidiaries” rather than “not ready.” Remove items that no longer connect to a meaningful decision.

The radar’s most valuable output is not a prediction. It is preparedness: a set of opportunities already decomposed, tested, and bounded when an enabling capability crosses the threshold. Competitors may see the same announcement. Your advantage is knowing precisely what to do next.

Common mistake: rewarding novelty

Novel signals will always compete for attention. Give priority to evidence that changes a decision, not developments that merely sound important. A quiet improvement in recovery, permissions, or cost predictability may unlock a product sooner than a dramatic model release.

The concrete outcome: a portfolio of executable bets

After the first cycle, the radar should contain few enough items to govern seriously. Every item should state a capability, its boundaries, the workflow it could reshape, the failure that matters, and the next experiment.

The deeper shift is methodological. Instead of asking what artificial intelligence will do, the team observes what has become reliable, identifies which assumption that change invalidates, and prepares a reversible move. Foresight becomes less theatrical and more useful: not certainty about the future, but a disciplined capacity to recognize it early.

This post was drafted with AI assistance and reviewed against our editorial policy before publication. Corrections are made at the source, on the page, with the date shown.

AI strategytechnology scoutingproduct discoverycapability mappinginnovation

From our own rounds

Measured on The Curator, from real sessions people played on this site — not a third-party dataset.

Rounds played here
169
Questions per round
1.7
Play a round and add to these numbers
Share this post

Rate this article

No ratings yet

Discussion

Comments are moderated. Read our editorial policy.