Qualified models

Known here, or qualified safely on your device.

Talavine publishes Pass or Fail for known exact API, subscription CLI, and local model routes. Model Scout reuses a website Pass instead of qualifying that route again. If it discovers a promising exact route that is absent here, it can earn a bounded device-local Pass or Fail while another healthy route keeps working. API, subscription-CLI, and local identities are independent. Learn how Model Scout uses the catalog →

Release catalog

Ready — at least one model passes.

Talavine uses a healthy available model automatically. Qualification is separate from current authentication, quota, installation, and health.

Provider Route Exact model Status Qualified Safe reason
Ollama Local gemma4:12b-mlx Pass Aug 20, 2026 Qualified for Talavine generative workloads; local availability and memory are checked on this device.
Ollama Local hf.co/unsloth/gemma-4-12B-it-GGUF:UD-Q4_K_XL Pass Aug 20, 2026 Qualified for Talavine generative workloads; local availability and memory are checked on this device.
Ollama Local qwen3.5:35b-a3b Pass Aug 20, 2026 Qualified for Talavine generative workloads; local availability and memory are checked on this device.
Ollama Local qwen3.8:27b Pass Aug 20, 2026 Qualified for Talavine generative workloads; local availability and memory are checked on this device.
ClaudeCode Subscription CLI claude-code:claude-fable-5 Pass Aug 20, 2026 Qualified through the Claude subscription CLI, but reserved for explicit or priority use because it consumes scarce Claude session capacity quickly.
ClaudeCode Subscription CLI claude-code:haiku Pass Aug 20, 2026 Qualified through the Claude subscription CLI; authentication, quota, and current CLI health are checked locally.
ClaudeCode Subscription CLI claude-code:opus Pass Aug 20, 2026 Qualified through the Claude subscription CLI; authentication, quota, and current CLI health are checked locally.
ClaudeCode Subscription CLI claude-code:sonnet Pass Aug 20, 2026 Qualified through the Claude subscription CLI; authentication, quota, and current CLI health are checked locally.
CodexCli Subscription CLI codex-cli:gpt-5.6-luna Pass Aug 20, 2026 Qualified through the ChatGPT subscription CLI; authentication, quota, and current CLI health are checked locally.
CodexCli Subscription CLI codex-cli:gpt-5.6-sol Pass Aug 20, 2026 Qualified through the ChatGPT subscription CLI; authentication, quota, and current CLI health are checked locally.
CodexCli Subscription CLI codex-cli:gpt-5.6-terra Pass Aug 20, 2026 Qualified through the ChatGPT subscription CLI; authentication, quota, and current CLI health are checked locally.
CodexCli Subscription CLI codex-cli:gpt-6-astra Pass Sep 3, 2026 Qualified through the ChatGPT subscription CLI; authentication, quota, and current CLI health are checked locally.
Advanced diagnostics — optional function benchmarks and proof details

Most people never need this section. Ordinary function measurements support engineering and troubleshooting; they cannot independently qualify a model for automatic use. Model Scout's dedicated bounded qualification run may use a representative suite to issue an exact device-local Pass or Fail.

Optional engineering diagnostics

Function-specific benchmark evidence.

These measurements help engineering compare implementations. They do not automatically add a model to the Pass/Fail catalog and they are not a normal publish blocker.

Function Hardware bucket Model Score Pass Latency Confidence Required settings
AI chat assistant generic OpenAI/gpt-5.5 0.940 67% 34.1s official-lab
AI chat assistant subscription-cli ClaudeCode/claude-code:opus frontier tier Not measured Not measured official-operator-clearance provider_kind=subscription-cli requires_subscription=True
Embeddings windows-arm64-cpu-48gb Ollama/qwen3-embedding:0.6b 0.938 100% 0.1s official-lab keep_alive=10m query_instruction=communication-retrieval-v1
Official proof

Benchmarks and operator clearances.

Benchmark rows contain aggregate calibration results only. Operator-clearance rows are explicitly labelled and contain no invented score, pass-rate, latency, or sample measurements. Neither kind includes raw messages, prompts, model responses, headers, extracted facts, embeddings, or personal data.

Message pipeline

Gate: score at least 0.88, pass rate at least 90%, no critical failures, and no parse failures.

Scope Evidence Hardware bucket Model Quality Pass Latency Gate Notes

AI chat assistant

Measured gate: score at least 0.88, pass rate at least 60%, no critical failures, and no parse failures. Exact operator clearances are shown separately without measurements.

Scope Evidence Hardware bucket Model Quality Pass Latency Gate Notes
Cloud Measured benchmark generic OpenAI/gpt-5.5 0.940 67% 34.1s Clears Official demo-profile assistant harness refreshed 2026-06-05 UTC on three source-backed scenarios: expense itemization, last-month receipts, and mail-thread intelligence. Aggregate metadata only; one strict relative-date clarification near miss.
Cloud Operator clearance subscription-cli ClaudeCode/claude-code:opus frontier tier Not measured Not measured Cleared Operator-cleared for the current agent-chat contract; this is not a benchmark measurement. The exact Claude Code subscription route still requires a reachable authenticated CLI, exact advertised model, available quota, and the strict tool-use envelope at runtime.
Cloud Operator clearance subscription-cli ClaudeCode/claude-code:sonnet frontier tier Not measured Not measured Cleared Operator-cleared for the current agent-chat contract; this is not a benchmark measurement. The exact Claude Code subscription route still requires a reachable authenticated CLI, exact advertised model, available quota, and the strict tool-use envelope at runtime.
Cloud Operator clearance subscription-cli ClaudeCode/claude-code:haiku balanced tier Not measured Not measured Cleared Operator-cleared for the current agent-chat contract; this is not a benchmark measurement. The exact Claude Code subscription route still requires a reachable authenticated CLI, exact advertised model, available quota, and the strict tool-use envelope at runtime.

Embeddings

Gate: score at least 0.90, pass rate at least 90%, no critical failures, and no parse failures.

Scope Evidence Hardware bucket Model Quality Pass Latency Gate Notes
Local Measured benchmark windows-arm64-cpu-48gb Ollama/qwen3-embedding:0.6b 0.938 100% 0.1s Clears Official embedding-retrieval-v1 corpus-bound measurement on Windows ARM64 Snapdragon X2 hardware through Ollama CPU execution; Recall@1 0.875, Recall@3 1.0, MRR 0.9375, mean 138.73 ms, p95 157.68 ms. Aggregate metadata only; Ollama did not use the NPU for this embedding route.
Community data

What users can safely submit.

Allowed

Hardware bucket, OS/runtime version, model tag, quantization, score, pass rate, failure counts, latency, and safe runtime settings.

Never collected

Raw emails, prompts, model responses, headers, extracted facts, contact names, or any user-derived message content.

Bundle API

Catalog and engineering endpoints.

The desktop app caches this catalog as last-known-good data. It uses an exact website Pass only for the matching provider, model, and route kind, while checking current route health locally. Website-unknown routes can receive a separate, expiring device-local qualification receipt; those private receipts are not uploaded by this endpoint. Function benchmark endpoints remain available for optional engineering work.

GET /api/model-quality GET /api/model-quality/routing-bundle GET /api/model-benchmarks/message-pipeline GET /api/model-benchmarks/agent-chat GET /api/model-benchmarks/text-embedding GET /api/model-benchmarks/spam-detection POST /api/model-benchmarks/message-pipeline/community POST /api/model-benchmarks/{functionType}/community 2026.10.09 expires Oct 11, 2026