Allowed
Hardware bucket, OS/runtime version, model tag, quantization, score, pass rate, failure counts, latency, and safe runtime settings.
Talavine publishes Pass or Fail for known exact API, subscription CLI, and local model routes. Model Scout reuses a website Pass instead of qualifying that route again. If it discovers a promising exact route that is absent here, it can earn a bounded device-local Pass or Fail while another healthy route keeps working. API, subscription-CLI, and local identities are independent. Learn how Model Scout uses the catalog →
Talavine uses a healthy available model automatically. Qualification is separate from current authentication, quota, installation, and health.
| Provider | Route | Exact model | Status | Qualified | Safe reason |
|---|---|---|---|---|---|
| Ollama | Local | gemma4:12b-mlx |
Pass | Aug 20, 2026 | Qualified for Talavine generative workloads; local availability and memory are checked on this device. |
| Ollama | Local | hf.co/unsloth/gemma-4-12B-it-GGUF:UD-Q4_K_XL |
Pass | Aug 20, 2026 | Qualified for Talavine generative workloads; local availability and memory are checked on this device. |
| Ollama | Local | qwen3.5:35b-a3b |
Pass | Aug 20, 2026 | Qualified for Talavine generative workloads; local availability and memory are checked on this device. |
| Ollama | Local | qwen3.8:27b |
Pass | Aug 20, 2026 | Qualified for Talavine generative workloads; local availability and memory are checked on this device. |
| ClaudeCode | Subscription CLI | claude-code:claude-fable-5 |
Pass | Aug 20, 2026 | Qualified through the Claude subscription CLI, but reserved for explicit or priority use because it consumes scarce Claude session capacity quickly. |
| ClaudeCode | Subscription CLI | claude-code:haiku |
Pass | Aug 20, 2026 | Qualified through the Claude subscription CLI; authentication, quota, and current CLI health are checked locally. |
| ClaudeCode | Subscription CLI | claude-code:opus |
Pass | Aug 20, 2026 | Qualified through the Claude subscription CLI; authentication, quota, and current CLI health are checked locally. |
| ClaudeCode | Subscription CLI | claude-code:sonnet |
Pass | Aug 20, 2026 | Qualified through the Claude subscription CLI; authentication, quota, and current CLI health are checked locally. |
| CodexCli | Subscription CLI | codex-cli:gpt-5.6-luna |
Pass | Aug 20, 2026 | Qualified through the ChatGPT subscription CLI; authentication, quota, and current CLI health are checked locally. |
| CodexCli | Subscription CLI | codex-cli:gpt-5.6-sol |
Pass | Aug 20, 2026 | Qualified through the ChatGPT subscription CLI; authentication, quota, and current CLI health are checked locally. |
| CodexCli | Subscription CLI | codex-cli:gpt-5.6-terra |
Pass | Aug 20, 2026 | Qualified through the ChatGPT subscription CLI; authentication, quota, and current CLI health are checked locally. |
| CodexCli | Subscription CLI | codex-cli:gpt-6-astra |
Pass | Sep 3, 2026 | Qualified through the ChatGPT subscription CLI; authentication, quota, and current CLI health are checked locally. |
Most people never need this section. Ordinary function measurements support engineering and troubleshooting; they cannot independently qualify a model for automatic use. Model Scout's dedicated bounded qualification run may use a representative suite to issue an exact device-local Pass or Fail.
These measurements help engineering compare implementations. They do not automatically add a model to the Pass/Fail catalog and they are not a normal publish blocker.
| Function | Hardware bucket | Model | Score | Pass | Latency | Confidence | Required settings |
|---|---|---|---|---|---|---|---|
| AI chat assistant | generic | OpenAI/gpt-5.5 |
0.940 | 67% | 34.1s | official-lab | |
| AI chat assistant | subscription-cli | ClaudeCode/claude-code:opus |
frontier tier | Not measured | Not measured | official-operator-clearance | provider_kind=subscription-cli requires_subscription=True |
| Embeddings | windows-arm64-cpu-48gb | Ollama/qwen3-embedding:0.6b |
0.938 | 100% | 0.1s | official-lab | keep_alive=10m query_instruction=communication-retrieval-v1 |
Benchmark rows contain aggregate calibration results only. Operator-clearance rows are explicitly labelled and contain no invented score, pass-rate, latency, or sample measurements. Neither kind includes raw messages, prompts, model responses, headers, extracted facts, embeddings, or personal data.
Gate: score at least 0.88, pass rate at least 90%, no critical failures, and no parse failures.
| Scope | Evidence | Hardware bucket | Model | Quality | Pass | Latency | Gate | Notes |
|---|
Measured gate: score at least 0.88, pass rate at least 60%, no critical failures, and no parse failures. Exact operator clearances are shown separately without measurements.
| Scope | Evidence | Hardware bucket | Model | Quality | Pass | Latency | Gate | Notes |
|---|---|---|---|---|---|---|---|---|
| Cloud | Measured benchmark | generic | OpenAI/gpt-5.5 |
0.940 | 67% | 34.1s | Clears | Official demo-profile assistant harness refreshed 2026-06-05 UTC on three source-backed scenarios: expense itemization, last-month receipts, and mail-thread intelligence. Aggregate metadata only; one strict relative-date clarification near miss. |
| Cloud | Operator clearance | subscription-cli | ClaudeCode/claude-code:opus |
frontier tier | Not measured | Not measured | Cleared | Operator-cleared for the current agent-chat contract; this is not a benchmark measurement. The exact Claude Code subscription route still requires a reachable authenticated CLI, exact advertised model, available quota, and the strict tool-use envelope at runtime. |
| Cloud | Operator clearance | subscription-cli | ClaudeCode/claude-code:sonnet |
frontier tier | Not measured | Not measured | Cleared | Operator-cleared for the current agent-chat contract; this is not a benchmark measurement. The exact Claude Code subscription route still requires a reachable authenticated CLI, exact advertised model, available quota, and the strict tool-use envelope at runtime. |
| Cloud | Operator clearance | subscription-cli | ClaudeCode/claude-code:haiku |
balanced tier | Not measured | Not measured | Cleared | Operator-cleared for the current agent-chat contract; this is not a benchmark measurement. The exact Claude Code subscription route still requires a reachable authenticated CLI, exact advertised model, available quota, and the strict tool-use envelope at runtime. |
Gate: score at least 0.90, pass rate at least 90%, no critical failures, and no parse failures.
| Scope | Evidence | Hardware bucket | Model | Quality | Pass | Latency | Gate | Notes |
|---|---|---|---|---|---|---|---|---|
| Local | Measured benchmark | windows-arm64-cpu-48gb | Ollama/qwen3-embedding:0.6b |
0.938 | 100% | 0.1s | Clears | Official embedding-retrieval-v1 corpus-bound measurement on Windows ARM64 Snapdragon X2 hardware through Ollama CPU execution; Recall@1 0.875, Recall@3 1.0, MRR 0.9375, mean 138.73 ms, p95 157.68 ms. Aggregate metadata only; Ollama did not use the NPU for this embedding route. |
Hardware bucket, OS/runtime version, model tag, quantization, score, pass rate, failure counts, latency, and safe runtime settings.
Raw emails, prompts, model responses, headers, extracted facts, contact names, or any user-derived message content.
The desktop app caches this catalog as last-known-good data. It uses an exact website Pass only for the matching provider, model, and route kind, while checking current route health locally. Website-unknown routes can receive a separate, expiring device-local qualification receipt; those private receipts are not uploaded by this endpoint. Function benchmark endpoints remain available for optional engineering work.