Technology deep dive

Model Scout: qualify once, reuse everywhere.

Talavine publishes a simple Pass or Fail for known exact API, subscription-CLI, or local model routes. Model Scout checks that catalog first, skips qualification when the route already passes, and can run one bounded background qualification when it discovers a promising route the website does not know yet. Current authentication, quota, installation, and health remain separate checks. Talavine uses a healthy available model automatically.

Simple qualification

Exact model route: Pass or Fail.

A website qualification proves model quality. It does not pretend an API key has quota, a CLI is logged in, or a local model is installed; those volatile checks remain local and independent.

One reviewed decision

Qualification is keyed by exact provider, exact model, and route kind. A Claude API route and Claude subscription-CLI route never share health, quota, or qualification state.

Current health stays separate

After a Pass, Model Scout checks only what can change on this machine: availability, authentication, quota, endpoint health, installation, execution contract, and local memory fit.

Bounded unknown-model scouting

A known Pass is reused without another qualification, and a known Fail is not retried automatically. An unknown promising model can be screened and qualified one at a time in the background, under cooldown, quota, and spend limits, while a healthy route keeps serving.

User pins stay sovereign LocalOnly enforced before any cloud advisor Hard requirements checked at one capability boundary
Selection inputs

Qualification first, live health second.

A failed or unavailable route does not make the setup look broken when another exact configured route passes and is healthy.

Source What it contributes
Qualified-model catalog The website's versioned Pass/Fail decision for known exact provider, model, and route-kind identities.
Device-local qualification A bounded Pass/Fail receipt for an exact website-unknown route on this installation. API, subscription CLI, and local identities never substitute for one another.
Current route health Independent API, subscription-CLI, and local checks for availability, authentication, quota, installation, and execution compatibility.
User choices Pinned provider/model choices remain exact and sovereign; their current Pass/Fail and health are shown clearly.
Advanced diagnostics Optional function benchmarks, recovery details, and runtime telemetry remain available for engineering and troubleshooting.
Proof that expires correctly

Same model, different machine? Different decision.

Capability verdicts are cached under a composite key: prompt schema hash, provider and model, local model digest, runtime version, quantization, and a hardware fingerprint. Change any of them and stale proof invalidates itself — but a daily rebuild doesn't.

JSON-contract probe

A known website Pass avoids another qualification. For an unknown candidate, Model Scout can run a cheap screen followed by one bounded representative suite; incomplete, unavailable, quota-limited, or interrupted attempts remain Unknown rather than becoming Fail.

Tool-use calibration

Optional engineering calibration can still prove native tool calling per model, schema, and runtime. Those function-specific results can refine ranking, but they do not replace the exact model-level Pass/Fail decision.

Serving health

Sometimes the model is fine and the serving layer is sick.

A local model that re-pays a 22-second prompt prefill on every call isn't a bad model — it's a cache-eviction problem. Scout's serving-health lane diagnoses the serving layer separately from model quality.

  1. Capture serving telemetry Prefill time, decode rate, and model-load events are recorded on every local call, alongside runtime introspection of what's actually loaded and how.
  2. Run deterministic detectors Named anomaly signatures — repeated full prefill, alternation eviction, warm-up lag, load thrash — fire on the telemetry, not on guesswork.
  3. Fire synthetic probes Same-prefix-repeat and alternation probes confirm a suspected pathology before anything is changed.
  4. Remediate and verify A two-tier remediation catalog applies fixes — an in-app prefix-cache primer stays silent; environment tuning that changes your machine asks first — then re-probes to confirm the fix took.
Local scheduling

One GPU, many callers, zero churn.

The local runtime manager serializes inference to one in-flight request, with a drain gate that lets user-waiting calls jump ahead of background work.

Priority drain gate

Background enrichment yields to your question. The queue drains interactive work first, so a batch job never makes the assistant feel slow.

Keep-alive lanes

Vision, embedding, and interactive workloads each get their own keep-alive policy, so the runtime stops paying model load/unload churn every time work alternates.

Observable

Hot-model-per-lane and switch counts are tracked, so "why was that slow?" has an answer in data instead of a shrug.

Routing policy

Advisors recommend. The governor decides. The guard verifies.

Pluggable, precedence-ordered advisors shape routes before the proof gate — each one defers cleanly under privacy or urgency constraints.

Budget pressure

As spend approaches the cap, work steps down to local or economy tiers before the cap is hit — a soft landing instead of a hard cutoff.

Quota step-off

When a Claude Code or Codex subscription lane is saturated, the advisor steps off to paid API headroom (or local) rather than queueing behind an exhausted plan.

Urgency & timing

User-waiting calls can justify a faster or better route than overnight background work — the same function routes differently at different moments.

Subscription routes

Your coding-agent subscription is a first-class model route.

Talavine detects, installs, and authenticates subscription-backed CLIs (Claude Code, Codex, Gemini CLI) from inside the app — cross-platform executable discovery across npm, volta, nvm, scoop, and desktop-app install locations, CPU-architecture mismatch detection, auth-state probing, and OAuth launch with API-credential scrubbing so the CLI session can't silently bill your API key instead of your subscription.

Already-paid capacity used first Newest-semver resolution Auth probed, not assumed
Keep reading

Where the proof gate bites hardest.

The message pipeline is Model Scout's strictest customer — every promotion runs through calibration gates and production replay.