Skip to content
Sign in with Google

Models and Usage

Compare Bkper AI models, capabilities, usage rates, monthly allowances, and usage visibility.

Bkper AI offers one model catalog and one included usage allowance. Choose by the output your application needs:

  • Generated language — use a language model for conversations, text, structured JSON, tools, images, or files. Bkper CLI Agent uses these models.
  • A bounded judgment — use a decision model, such as Jev, when code needs a noul probability, one choice from supplied options, or an ordered score.

Decision models are not conversational. Apps and compatible clients call them through the Decisions API, and Bkper CLI Agent can ask them for you.

This page does not apply when you connect Bkper to ChatGPT or Claude and continue using the models provided by those products.

Models

Public IDs identify stable Bkper model families. Bkper can upgrade the concrete provider revision behind a family without changing its public ID, saved client configuration, or usage-report identity.

ModelWhat it offers
Claude Sonnet claude-sonnet
Language · CLI + Responses API
Fast, capable multimodal model for coding, reasoning, and agent workflows.
Reasoning: low, medium, high
Limits: 300k context · 64k output
DeepSeek Flash deepseek-flash
Language · CLI + Responses API
Fast, economical multimodal reasoning, coding, and long-context work.
Reasoning: high, max
Limits: 300k context · 64k output
Gemini Flash-Lite gemini-flash-lite
Language · CLI + Responses API
Economical multimodal model for fast, high-volume tasks and tool use.
Reasoning: minimal, low, medium, high
Limits: 1,048,576 context · 65,536 output
Gemini Flash gemini-flash
Language · CLI + Responses API
Fast multimodal reasoning and tool use with balanced usage rates.
Reasoning: low, medium
Limits: 1,048,576 context · 65,536 output
GPT Luna gpt-luna
Language · CLI + Responses API
Cost-efficient GPT model for fast, high-volume workloads.
Reasoning: high, xhigh, max
Limits: 272k context · 64k output
GPT Sol gpt-sol
Language · CLI + Responses API
High-capability GPT model for complex reasoning, coding, and demanding agent workflows.
Reasoning: low, medium, high
Limits: 272k context · 64k output
Grok grok
Language · CLI + Responses API
General-purpose model for chat, coding, and agentic tool use.
Reasoning: low, medium
Limits: 200k context · 32k output
GLM Flash glm-flash
Language · CLI + Responses API
Fast, economical multimodal reasoning, coding, and long-context work.
Reasoning: low, high, max
Limits: 1,048,576 context · 64k output
Jev jev
Decision · Decisions API
Fast typed judgments for noul, choice, and score questions.
Question types: noul, choice, score
Limits: 65,536 context · 32,768 state + question

Bkper CLI loads the current context window, output limit, and reasoning profile for language models from the live catalog. Compatible clients can request lower output budgets and any listed reasoning effort. Anthropic safeguards can refuse Claude Sonnet requests; a tool call is not guaranteed.

Jev evaluates one shared text or JSON state against independent questions in parallel. Its responses report the concrete Jev revision in model; the catalog and usage history keep the stable jev ID. They are complete JSON and do not stream or create conversation state. See Decision Models for requests, responses, and decision patterns.

How usage works

Bkper AI is included with eligible plans and controlled through one monthly allowance:

  • The Bkper CLI Agent connects with your Bkper login; no separate provider setup is required.
  • Model usage reduces the included allowance and is not billed separately.
  • New requests stop when the recorded allowance is exhausted; there are no automatic paid overages.
  • External providers remain available where supported and do not consume the Bkper AI allowance.

Requests already in flight may settle after the allowance check, so recorded usage can slightly exceed the limit under concurrency. Bkper does not bill that difference as an AI overage.

Usage rates

Usage rates reduce the included monthly allowance. They are not billed separately by Bkper.

USD of included usage per one million tokens

ModelInputCache readCache writeOutput
Claude Sonnet$3.00$0.30$3.80$15.00
DeepSeek Flash$0.45$0.009$0.00$1.80
Gemini Flash-Lite$0.45$0.045$0.00$3.80
Gemini Flash$1.10$0.11$0.00$5.60
GPT Luna$0.15$0.015$0.19$0.75
GPT Sol$3.00$0.15$3.80$15.00
Grok$3.00$0.75$0.00$9.00
GLM Flash$0.23$0.045$0.00$0.75
Jev (decision)$0.063——$0.00

Input means tokens sent without a cache match. Cache read means reused input already stored by the provider. Cache write means input added to a provider cache. Output includes generated response and reasoning tokens reported by the provider. Jev is input-only, so cache rates do not apply and returned typed answers add no output-token rate.

Claude Sonnet refusals are reported separately from successful responses. Documented unbilled pre-output Anthropic refusals retain reported token counts but consume no allowance; other refusals may consume allowance.

Monthly allowance

For paid plans, the monthly AI allowance equals the normalized monthly software subscription value:

  • Monthly plans use the recurring monthly software subscription value.
  • Annual plans divide the recurring annual software subscription value by 12.

Only recurring software subscription value counts. Professional services, implementation, consulting, taxes, one-time charges, credits, refunds, and prorations do not increase the allowance. Free users receive a separately configured trial allowance.

The allowance resets monthly and unused value does not roll over. It is an inference entitlement—not cash, refund value, or transferable credit. The authenticated Bkper AI usage dashboard shows your exact current allowance.

Individual and pooled usage

Allowance scope follows the subscription:

  • Free and Standard usage is assigned to the individual user.
  • Business usage is pooled across the subscribed domain.
  • Everyone sharing a pooled allowance reduces the same monthly total.

A pooled allowance does not make every user’s request history visible to everyone. Visibility depends on the viewer’s billing role.

Usage visibility and privacy

The Bkper AI usage dashboard separates allowance visibility from request attribution:

  • Regular users see the shared allowance remaining and their own requests and usage.
  • The billing or subscription administrator sees domain-wide usage attributed by user, AI model, and app or source.
  • The dashboard does not expose prompts or responses.

Usage value is an estimate based on the published rates above. It shows how much of the included allowance a request consumed; it is not a separate Bkper charge.

When the limit is reached

Bkper AI blocks new allowance-backed requests once the recorded monthly allowance is exhausted. There are no automatic paid Bkper AI overages at launch.

Bkper AI is not a lock-in: where supported, you can connect an external model provider at any time. External subscriptions, API keys, charges, privacy terms, and limits are governed by that provider and do not use the included Bkper AI allowance.

How Bkper selects language models

We build Bkper with the Bkper CLI Agent and use it every day. We test many language models through real work and include only those that consistently work well for us within our cost and control constraints. The catalog is a practical, opinionated shortlist—not a directory of every available model.

Bkper prioritizes strong results at controlled cost — the efficient frontier of capability per dollar — rather than pursuing the highest benchmark score at any price. Selection also considers:

  • results and reliability in daily agent workflows;
  • effective model capabilities and tool use;
  • observed usage cost;
  • model capabilities and controls;
  • public benchmarks.

Explore the live DeepSWE leaderboard.

DeepSWE measures long-horizon software-engineering work. It is one input into model selection, not a measure of accounting accuracy or a guarantee of performance in Bkper workflows.

Model creators

ModelTypeCreator
Claude SonnetLanguageAnthropic
DeepSeek FlashLanguageDeepSeek
Gemini Flash-LiteLanguageGoogle
Gemini FlashLanguageGoogle
GPT LunaLanguageOpenAI
GPT SolLanguageOpenAI
GrokLanguagexAI
GLM FlashLanguageZ.ai
JevDecisionTypeSafe

Sources

Last synchronized: 2026-10-03

Models.dev provider assets are provided under the MIT License. Model creator names and logos remain trademarks of their respective owners.

Next steps