KORA Doctor flags LLM calls your AI agent may never have needed

LLM costs add up quickly in agent runs. AUDR records the spend, but KORA Doctor asks whether all that inference was actually necessary. The CLI analyzes AUDR JSON/JSONL and flags calls worth removing, caching, replacing with deterministic logic, or moving to a cheaper model. No dashboard, no account, no hosted service.

Install and audit in one command

Install directly from GitHub:

pipx install git+https://github.com/Krako-Labs/kora-doctor.git

Then audit an AUDR file:

kora-doctor audit audr.jsonl

From a code clone:

python3 -m kora_doctor audit samples/inefficient_agent.jsonl

Machine-readable output:

kora-doctor audit audr.jsonl --json

Audit results show what can be optimized

Running the included inefficient trace produces this summary:

Observed: 11 records · 2 runs · 11 model calls · 0 tool calls
Observed cost:                      $0.1050
Potentially avoidable:              $0.0936  (89%)
Estimated optimized cost:           $0.0114

Candidates flagged for review:

  • Duplicate/repeated calls: 5
  • Cache/reuse candidates: 5
  • Deterministic candidates: 11
  • Smaller-model candidates: 11
  • Orchestration overhead: 5

How KORA Doctor identifies candidates

KORA Doctor v0 uses simple heuristics, not proofs. The goal is to narrow long traces down to calls a developer should inspect first.

  • Duplicate/repeated inference — the same model/resource and usage signature repeating inside one run
  • Cache/reuse candidates — the same signature appearing across multiple runs
  • Deterministic candidates — model calls whose AUDR run metadata looks like classification, routing, validation, extraction, formatting, parsing, or normalization
  • Smaller-model candidates — short calls on high-end models with no reported reasoning tokens
  • Suspicious orchestration overhead — agent runs with many model calls where later calls deserve inspection

How the tool works under the hood

KORA Doctor checks core fields needed for analysis. It is not a replacement for the official AUDR JSON Schema conformance validator.

Key principles:

  • Observed — values directly present in the trace
  • Estimated — derived savings calculations
  • Candidate — optimization opportunities
  • Confidence — heuristic strength
  • Insufficient evidence — where the trace cannot support a stronger conclusion

Savings estimates and limitations

Savings are only calculated when cost.total_cost is present. Multiple currencies are never silently converted.

Each candidate type has a scenario ratio:

  • Duplicate/repeated call: 100% of that call's observed cost
  • Cache/reuse candidate: 70%
  • Deterministic candidate: 80%
  • Smaller-model candidate: 50%
  • Orchestration-overhead candidate: 50%

If one call matches several rules, KORA Doctor uses only the largest ratio for that call; it never stacks savings estimates. These defaults are intentionally easy to inspect and change as real traces arrive.

The dollar estimate is a scenario estimate attached to each candidate, not a measured future bill. AUDR v1.0.0 deliberately records usage/cost telemetry without prompt content or secrets. That is good for security, but it also means an AUDR record alone usually cannot prove that two model calls were semantically identical or that a task could definitely have been deterministic. KORA Doctor therefore uses observed values and heuristic inference only.

File formats and workflow

KORA Doctor currently targets AUDR v1.0.0 and accepts:

  • One AUDR JSON object
  • A JSON array of AUDR objects
  • JSONL with one AUDR object per line

Sample audits:

python3 -m kora_doctor audit samples/simple.jsonl
python3 -m kora_doctor audit samples/multi_step.jsonl
python3 -m kora_doctor audit samples/inefficient_agent.jsonl

The samples are synthetic AUDR-compatible traces created for KORA Doctor. The inefficient trace is intentionally constructed to trigger multiple heuristics.

No runtime dependencies

KORA Doctor has no runtime dependencies. Install with:

python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install -e .
python3 -m unittest discover -s tests -v

The demo script:

./scripts/demo.sh

Notes from the demo:

  • No prompt/input fingerprints means duplicate and cache findings are heuristic
  • Deterministic candidates are inferred from run names and labels only
  • Smaller-model recommendations do not benchmark output quality
  • Estimated savings are scenario estimates, not guaranteed savings
  • KORA Doctor does not modify your agent or automatically reroute traffic in v0

These are deliberate v0 constraints. If real traces show that an extra signal is necessary, the smallest useful one will be added.

KORA Doctor is standalone. It does not require the full KORA runtime.

Future direction

Today the flow is: AUDR trace → KORA Doctor → diagnose. If users pull for it later: AUDR trace → diagnose → recommend → optimize automatically with KORA. Real AUDR traces, false positives, and missed optimization opportunities are the highest-value feedback. See CONTRIBUTING.md or open an issue.

License: Apache-2.0.