One reviewed decision
Qualification is keyed by exact provider, exact model, and route kind. A Claude API route and Claude subscription-CLI route never share health, quota, or qualification state.
Talavine publishes a simple Pass or Fail for known exact API, subscription-CLI, or local model routes. Model Scout checks that catalog first, skips qualification when the route already passes, and can run one bounded background qualification when it discovers a promising route the website does not know yet. Current authentication, quota, installation, and health remain separate checks. Talavine uses a healthy available model automatically.
A website qualification proves model quality. It does not pretend an API key has quota, a CLI is logged in, or a local model is installed; those volatile checks remain local and independent.
Qualification is keyed by exact provider, exact model, and route kind. A Claude API route and Claude subscription-CLI route never share health, quota, or qualification state.
After a Pass, Model Scout checks only what can change on this machine: availability, authentication, quota, endpoint health, installation, execution contract, and local memory fit.
A known Pass is reused without another qualification, and a known Fail is not retried automatically. An unknown promising model can be screened and qualified one at a time in the background, under cooldown, quota, and spend limits, while a healthy route keeps serving.
A failed or unavailable route does not make the setup look broken when another exact configured route passes and is healthy.
| Source | What it contributes |
|---|---|
| Qualified-model catalog | The website's versioned Pass/Fail decision for known exact provider, model, and route-kind identities. |
| Device-local qualification | A bounded Pass/Fail receipt for an exact website-unknown route on this installation. API, subscription CLI, and local identities never substitute for one another. |
| Current route health | Independent API, subscription-CLI, and local checks for availability, authentication, quota, installation, and execution compatibility. |
| User choices | Pinned provider/model choices remain exact and sovereign; their current Pass/Fail and health are shown clearly. |
| Advanced diagnostics | Optional function benchmarks, recovery details, and runtime telemetry remain available for engineering and troubleshooting. |
Capability verdicts are cached under a composite key: prompt schema hash, provider and model, local model digest, runtime version, quantization, and a hardware fingerprint. Change any of them and stale proof invalidates itself — but a daily rebuild doesn't.
A known website Pass avoids another qualification. For an unknown candidate, Model Scout can run a cheap screen followed by one bounded representative suite; incomplete, unavailable, quota-limited, or interrupted attempts remain Unknown rather than becoming Fail.
Optional engineering calibration can still prove native tool calling per model, schema, and runtime. Those function-specific results can refine ranking, but they do not replace the exact model-level Pass/Fail decision.
A local model that re-pays a 22-second prompt prefill on every call isn't a bad model — it's a cache-eviction problem. Scout's serving-health lane diagnoses the serving layer separately from model quality.
The local runtime manager serializes inference to one in-flight request, with a drain gate that lets user-waiting calls jump ahead of background work.
Background enrichment yields to your question. The queue drains interactive work first, so a batch job never makes the assistant feel slow.
Vision, embedding, and interactive workloads each get their own keep-alive policy, so the runtime stops paying model load/unload churn every time work alternates.
Hot-model-per-lane and switch counts are tracked, so "why was that slow?" has an answer in data instead of a shrug.
Pluggable, precedence-ordered advisors shape routes before the proof gate — each one defers cleanly under privacy or urgency constraints.
As spend approaches the cap, work steps down to local or economy tiers before the cap is hit — a soft landing instead of a hard cutoff.
When a Claude Code or Codex subscription lane is saturated, the advisor steps off to paid API headroom (or local) rather than queueing behind an exhausted plan.
User-waiting calls can justify a faster or better route than overnight background work — the same function routes differently at different moments.
Talavine detects, installs, and authenticates subscription-backed CLIs (Claude Code, Codex, Gemini CLI) from inside the app — cross-platform executable discovery across npm, volta, nvm, scoop, and desktop-app install locations, CPU-architecture mismatch detection, auth-state probing, and OAuth launch with API-credential scrubbing so the CLI session can't silently bill your API key instead of your subscription.
The message pipeline is Model Scout's strictest customer — every promotion runs through calibration gates and production replay.