Models and Usage
Compare Bkper AI models, capabilities, usage rates, monthly allowances, and usage visibility.
Bkper AI offers one model catalog and one included usage allowance. Choose by the output your application needs:
- Generated language — use a language model for conversations, text, structured JSON, tools, images, or files. Bkper CLI Agent uses these models.
- A bounded judgment — use a decision model, such as Jev, when code needs a
noulprobability, one choice from supplied options, or an ordered score.
Decision models are not conversational. Apps and compatible clients call them through the Decisions API, and Bkper CLI Agent can ask them for you.
This page does not apply when you connect Bkper to ChatGPT or Claude and continue using the models provided by those products.
Models
Public IDs identify stable Bkper model families. Bkper can upgrade the concrete provider revision behind a family without changing its public ID, saved client configuration, or usage-report identity.
| Model | What it offers |
|---|---|
Claude Sonnet claude-sonnetLanguage · CLI + Responses API | Fast, capable multimodal model for coding, reasoning, and agent workflows. Reasoning: low, medium, highLimits: 300k context · 64k output |
DeepSeek Flash deepseek-flashLanguage · CLI + Responses API | Fast, economical multimodal reasoning, coding, and long-context work. Reasoning: high, maxLimits: 300k context · 64k output |
Gemini Flash-Lite gemini-flash-liteLanguage · CLI + Responses API | Economical multimodal model for fast, high-volume tasks and tool use. Reasoning: minimal, low, medium, highLimits: 1,048,576 context · 65,536 output |
Gemini Flash gemini-flashLanguage · CLI + Responses API | Fast multimodal reasoning and tool use with balanced usage rates. Reasoning: low, mediumLimits: 1,048,576 context · 65,536 output |
GPT Luna gpt-lunaLanguage · CLI + Responses API | Cost-efficient GPT model for fast, high-volume workloads. Reasoning: high, xhigh, maxLimits: 272k context · 64k output |
GPT Sol gpt-solLanguage · CLI + Responses API | High-capability GPT model for complex reasoning, coding, and demanding agent workflows. Reasoning: low, medium, highLimits: 272k context · 64k output |
Grok grokLanguage · CLI + Responses API | General-purpose model for chat, coding, and agentic tool use. Reasoning: low, mediumLimits: 200k context · 32k output |
GLM Flash glm-flashLanguage · CLI + Responses API | Fast, economical multimodal reasoning, coding, and long-context work. Reasoning: low, high, maxLimits: 1,048,576 context · 64k output |
Jev jevDecision · Decisions API | Fast typed judgments for noul, choice, and score questions. Question types: noul, choice, scoreLimits: 65,536 context · 32,768 state + question |
Bkper CLI loads the current context window, output limit, and reasoning profile for language models from the live catalog. Compatible clients can request lower output budgets and any listed reasoning effort. Anthropic safeguards can refuse Claude Sonnet requests; a tool call is not guaranteed.
Jev evaluates one shared text or JSON state against independent questions in parallel. Its responses report the concrete Jev revision in model; the catalog and usage history keep the stable jev ID. They are complete JSON and do not stream or create conversation state. See Decision Models for requests, responses, and decision patterns.
How usage works
Bkper AI is included with eligible plans and controlled through one monthly allowance:
- The Bkper CLI Agent connects with your Bkper login; no separate provider setup is required.
- Model usage reduces the included allowance and is not billed separately.
- New requests stop when the recorded allowance is exhausted; there are no automatic paid overages.
- External providers remain available where supported and do not consume the Bkper AI allowance.
Requests already in flight may settle after the allowance check, so recorded usage can slightly exceed the limit under concurrency. Bkper does not bill that difference as an AI overage.
Usage rates
Usage rates reduce the included monthly allowance. They are not billed separately by Bkper.
USD of included usage per one million tokens
| Model | Input | Cache read | Cache write | Output |
|---|---|---|---|---|
| Claude Sonnet | $3.00 | $0.30 | $3.80 | $15.00 |
| DeepSeek Flash | $0.45 | $0.009 | $0.00 | $1.80 |
| Gemini Flash-Lite | $0.45 | $0.045 | $0.00 | $3.80 |
| Gemini Flash | $1.10 | $0.11 | $0.00 | $5.60 |
| GPT Luna | $0.15 | $0.015 | $0.19 | $0.75 |
| GPT Sol | $3.00 | $0.15 | $3.80 | $15.00 |
| Grok | $3.00 | $0.75 | $0.00 | $9.00 |
| GLM Flash | $0.23 | $0.045 | $0.00 | $0.75 |
| Jev (decision) | $0.063 | — | — | $0.00 |
Input means tokens sent without a cache match. Cache read means reused input already stored by the provider. Cache write means input added to a provider cache. Output includes generated response and reasoning tokens reported by the provider. Jev is input-only, so cache rates do not apply and returned typed answers add no output-token rate.
Claude Sonnet refusals are reported separately from successful responses. Documented unbilled pre-output Anthropic refusals retain reported token counts but consume no allowance; other refusals may consume allowance.
Monthly allowance
For paid plans, the monthly AI allowance equals the normalized monthly software subscription value:
- Monthly plans use the recurring monthly software subscription value.
- Annual plans divide the recurring annual software subscription value by 12.
Only recurring software subscription value counts. Professional services, implementation, consulting, taxes, one-time charges, credits, refunds, and prorations do not increase the allowance. Free users receive a separately configured trial allowance.
The allowance resets monthly and unused value does not roll over. It is an inference entitlement—not cash, refund value, or transferable credit. The authenticated Bkper AI usage dashboard shows your exact current allowance.
Individual and pooled usage
Allowance scope follows the subscription:
- Free and Standard usage is assigned to the individual user.
- Business usage is pooled across the subscribed domain.
- Everyone sharing a pooled allowance reduces the same monthly total.
A pooled allowance does not make every user’s request history visible to everyone. Visibility depends on the viewer’s billing role.
Usage visibility and privacy
The Bkper AI usage dashboard separates allowance visibility from request attribution:
- Regular users see the shared allowance remaining and their own requests and usage.
- The billing or subscription administrator sees domain-wide usage attributed by user, AI model, and app or source.
- The dashboard does not expose prompts or responses.
Usage value is an estimate based on the published rates above. It shows how much of the included allowance a request consumed; it is not a separate Bkper charge.
When the limit is reached
Bkper AI blocks new allowance-backed requests once the recorded monthly allowance is exhausted. There are no automatic paid Bkper AI overages at launch.
Bkper AI is not a lock-in: where supported, you can connect an external model provider at any time. External subscriptions, API keys, charges, privacy terms, and limits are governed by that provider and do not use the included Bkper AI allowance.
How Bkper selects language models
We build Bkper with the Bkper CLI Agent and use it every day. We test many language models through real work and include only those that consistently work well for us within our cost and control constraints. The catalog is a practical, opinionated shortlist—not a directory of every available model.
Bkper prioritizes strong results at controlled cost — the efficient frontier of capability per dollar — rather than pursuing the highest benchmark score at any price. Selection also considers:
- results and reliability in daily agent workflows;
- effective model capabilities and tool use;
- observed usage cost;
- model capabilities and controls;
- public benchmarks.
Explore the live DeepSWE leaderboard.
DeepSWE measures long-horizon software-engineering work. It is one input into model selection, not a measure of accounting accuracy or a guarantee of performance in Bkper workflows.
Model creators
| Model | Type | Creator |
|---|---|---|
| Claude Sonnet | Language | Anthropic |
| DeepSeek Flash | Language | DeepSeek |
| Gemini Flash-Lite | Language | |
| Gemini Flash | Language | |
| GPT Luna | Language | OpenAI |
| GPT Sol | Language | OpenAI |
| Grok | Language | xAI |
| GLM Flash | Language | Z.ai |
| Jev | Decision | TypeSafe |
Sources
Last synchronized: 2026-10-03
Models.dev provider assets are provided under the MIT License. Model creator names and logos remain trademarks of their respective owners.