Typed Evaluations
Ask an evaluation model bounded questions about a state and get typed, calibrated answers your code can act on.
A typed evaluation asks a model bounded questions about a state. Each answer is typed: the probability that a statement is true, one option from a set you define, or a level on a scale you define. The model writes no text. Your code reads the answer and decides what happens next.
POST /v1/evaluations serves evaluation models: models trained for calibrated decisions. Jev, from TypeSafe, is the first. The interface and its concepts come from TypeSafe’s System One API, and their documentation is a good companion to this page.
Why evaluations fit operations
Language models are trained to produce text people prefer. That suits conversation, code, and reasoning. But a fluent answer is not necessarily a reliable one, and turning text into decisions means parsing it. Evaluation models are trained for calibrated decisions: across many answers, those given a probability of 0.8 should be right about 80% of the time.
Accounting operations — categorizing bank lines, spotting duplicates, flagging entries for review — repeat the same kind of decision thousands of times, and each one must hold up to review. Evaluations give each decision three properties:
- Bounded — the answer is always one of the options you defined.
- Calibrated — probability and confidence tell your code when to act and when to ask a person.
- Auditable — thresholds live in your code, and every answer reports the model revision that produced it.
| When your code needs | Use |
|---|---|
| A yes/no, one of known options, or a level on a scale | A typed evaluation |
| Text, explanations, code, tool calls, or a custom JSON object | A language model |
TypeSafe’s AI primer explains how calibrated models differ from chat models in training.
How an evaluation works
A request has one state and one or more questions. The state is what the model judges: a string, JSON object, or array. Prefer an object with descriptive field names. Evaluation models currently read text only, so convert images and files to text first.
The model evaluates each question independently, in parallel, against the same state, and returns one typed answer per question. There are three question types:
noul— Is this true? Optionalcriteriadescribe what yes and no mean. Returnsnoul, the probability of yes from 0 to 1.choice— Which one of these?criteriamaps 1–255 options to descriptions, ornull. Returns thechoicewithprobabilitiesandconfidence.score— Which level on this scale?criterialists 2–10 levels, lowest first. Returns a weightedscorewithlegend,probabilities, andconfidence.
Ask one snap judgment per question — something a knowledgeable person decides in seconds. Split bigger judgments into several questions and combine the answers in code. Keep arithmetic, balances, and date comparisons in code too, and ask the model only the part that needs judgment. See TypeSafe’s Primitives and State.
Send a request
For an incoming bank line, this request asks three questions at once: which Account it belongs to, whether it is already recorded, and how urgently it needs review.
curl --fail-with-body https://ai.bkper.app/v1/evaluations \ -H "Authorization: Bearer ${BKPER_TOKEN}" \ -H "Content-Type: application/json" \ -H "bkper-ai-source: my-app" \ --data '{ "model": "jev", "state": { "bank_line": { "date": "2025-03-11", "description": "AMAZON MKTPL*2K4LM", "amount": 89.90, "direction": "money out of Checking" }, "existing_transaction": { "date": "2025-03-10", "amount": 89.90, "from": "Checking", "to": "Office Supplies", "description": "Printer toner, Amazon order 114-2K4LM", "origin": "manual entry" } }, "questions": { "account": { "type": "choice", "instructions": "Which Account should receive the money that left Checking in `bank_line`?", "criteria": { "Cloud Hosting": "Servers, cloud infrastructure, storage, and hosting providers", "Software Subscriptions": "SaaS tools and per-seat software licenses", "Office Supplies": "Physical supplies and equipment for the office", "Travel": "Flights, hotels, ground transport, and meals while traveling" } }, "already_recorded": { "type": "noul", "instructions": "Does `existing_transaction` already record the movement in `bank_line`?" }, "review_priority": { "type": "score", "instructions": "How urgently should a bookkeeper review `bank_line`?", "criteria": [ "Routine: consistent with normal activity", "Check: unusual but plausibly legitimate", "Urgent: large, unapproved, or suspicious" ] } } }'model— any model withtype: "evaluation"inGET /v1/models. Each entry lists itsquestion_types,context_window, andmax_state_question_tokens.- Question IDs such as
accountare yours. Answers come back under the same IDs, but IDs are not sent to the model, so eachinstructionsmust stand on its own. - Refer to state fields by name in backticks, as in
`bank_line`. - Each question accepts only
type,instructions, andcriteria. Instructions, descriptions, and levels can be strings or structured JSON.
Read the response
A real response to the request above:
{ "model": "jev-1.13.0", "answers": { "account": { "type": "choice", "choice": "Office Supplies", "confidence": 1, "probabilities": { "Office Supplies": 1, "Cloud Hosting": 0, "Travel": 0, "Software Subscriptions": 0 } }, "already_recorded": { "type": "noul", "noul": 0.89 }, "review_priority": { "type": "score", "score": 0.09, "confidence": 0.86, "legend": { "0": "Routine: consistent with normal activity", "1": "Check: unusual but plausibly legitimate", "2": "Urgent: large, unapproved, or suspicious" }, "probabilities": { "0": 0.91, "1": 0.09, "2": 0 } } }, "usage": { "input_tokens": 607, "output_tokens": 86 }}Bkper validates every answer before returning it, so your code can rely on these guarantees:
answershas exactly one entry per question ID, and each answer’stypematches its question.choiceis always one of yourcriteriakeys.scoreis the probability-weighted level, from0to the number of levels minus one. It can fall between levels.probabilitieshas one entry per option or level, andlegendmaps level indexes back to your text. Read both by key: order is not guaranteed, and values can be exactly0or1.confidence, on choice and score answers, runs from 0 to 1 and is higher when probability concentrates on one answer. Noul answers have noconfidence; the probability itself is the signal. See TypeSafe’s Confidence.modelis the concrete revision that answered, such asjev-1.13.0.
Usage counts against your allowance at the model’s usage rates.
Context changes confidence
Ask about the same bank line without existing_transaction:
{ "model": "jev", "state": { "bank_line": { "date": "2025-03-11", "description": "AMAZON MKTPL*2K4LM", "amount": 89.9, "direction": "money out of Checking" } }, "questions": { "account": { "type": "choice", "instructions": "Which Account should receive the money that left Checking in `bank_line`?", "criteria": { "Cloud Hosting": "Servers, cloud infrastructure, storage, and hosting providers", "Software Subscriptions": "SaaS tools and per-seat software licenses", "Office Supplies": "Physical supplies and equipment for the office", "Travel": "Flights, hotels, ground transport, and meals while traveling" } } }}{ "model": "jev-1.13.0", "answers": { "account": { "type": "choice", "choice": "Office Supplies", "confidence": 0.75, "probabilities": { "Software Subscriptions": 0.14, "Cloud Hosting": 0.02, "Office Supplies": 0.81, "Travel": 0.02 } } }, "usage": { "input_tokens": 440, "output_tokens": 52 }}The top option is the same, but confidence drops from 1 to 0.75: the model reports that it is less sure.
Give the model the context a person would need, such as nearby Transactions, how an Account is usually described, or Book and Account properties. Select that context in code. For example, find Transactions with the same amount inside a date window, then ask the model only whether the descriptions describe the same movement.
Turn answers into decisions
An answer is a judgment, not permission to change a Book. Your code owns the thresholds and the actions, and sends uncertain cases to a person. This is TypeSafe’s confidence-gated routing pattern:
// Keep questions and thresholds in one reviewable place.const DUPLICATE_THRESHOLD = 0.8;const REVIEW_SCORE_THRESHOLD = 1.5;const AUTO_POST_CONFIDENCE = 0.9;
interface BankLineAnswers { account: { type: 'choice'; choice: string; confidence: number }; already_recorded: { type: 'noul'; noul: number }; review_priority: { type: 'score'; score: number; confidence: number };}
type BankLineDecision = | { action: 'check-duplicate' } | { action: 'review'; suggestedAccount: string } | { action: 'post'; toAccount: string; confidence: number };
export function decideBankLine(answers: BankLineAnswers): BankLineDecision { if (answers.already_recorded.noul >= DUPLICATE_THRESHOLD) { return { action: 'check-duplicate' }; } if ( answers.review_priority.score >= REVIEW_SCORE_THRESHOLD || answers.account.confidence < AUTO_POST_CONFIDENCE ) { return { action: 'review', suggestedAccount: answers.account.choice }; } return { action: 'post', toAccount: answers.account.choice, confidence: answers.account.confidence, };}The first response returns check-duplicate: the movement is probably already recorded. The bank line on its own, at 0.75 confidence, goes to review.
When your code does post, it creates a normal Transaction from Checking to the chosen Account, so the Book stays zero-sum whatever the model answered. The risks to control are a wrong Account and a duplicate movement, which is why those checks come before posting. Store model and confidence as Transaction properties for the audit trail.
Start with conservative thresholds, test them against your own records, and re-check them when the returned model changes. For more designs, see TypeSafe’s Patterns and Cookbooks, and the open-source Merge Duplicates app, which suggests duplicate pairs for human review.
Call from a Bkper app
In a Bkper Platform app, call https://ai.bkper.app/v1/evaluations from the Worker, in an authenticated /api/* route or /events handler. Do not add an Authorization header. Platform outbound adds authorization and app attribution. Never read, forward, or store the user’s token.
Outside the platform — scripts, servers, local tests — send your own Bkper access token as a bearer token.
Errors and retries
Errors use the same envelope as every Bkper AI endpoint. Branch on error.code, not only on the HTTP status:
{ "error": { "message": "Unknown field: questions.route.threshold.", "type": "invalid_request_error", "param": "questions.route.threshold", "code": "invalid_request" }}| Status | error.code | What to do |
|---|---|---|
400 | invalid_request, missing_model, unsupported_model, invalid_state, invalid_questions, invalid_question, unsupported_question_type | Fix the field named in error.param. Do not retry unchanged. |
400 | provider_rejected | The model rejected the request. Check its size against the model’s limits. |
401–403 | unauthorized, billing_overdue, entitlement_unavailable | Fix authentication or account access. Do not retry. |
429 | usage_limit_exceeded | The monthly allowance is exhausted. Do not retry. |
429 | provider_rate_limited | Retry with backoff. Honor retry-after when present. |
503 | provider_overloaded, usage_unavailable | Retry with backoff. Honor retry-after when present. |
502 | provider_error, provider_rejected | Retry a limited number of times, then fail. |
Caching
Bkper caches evaluation results for 15 minutes per user. An identical request in that window returns the same answers without calling the model or consuming allowance, and repeats the original usage. Aliases of the same model share one cache entry. To get a fresh evaluation, change the request or wait. See Privacy, retention, and caching.
Common mistakes
- Sending Responses fields. The body accepts only
model,state, andquestions.input,stream,store,temperature, andmetadataare rejected. - Putting the question in the ID. The model never sees IDs. Write the full question in
instructions. - Asking for analysis. Ask one snap judgment per question and combine answers in code.
- Asking the model to calculate. Match amounts, compute date windows, and check balances in code.
- Treating
noulas a boolean orscoreas an integer. Both are continuous. Compare them with thresholds. - Expecting an explanation. Evaluations return no text. Use a language model when you need one.
Jev by TypeSafe
Jev is TypeSafe’s flagship model and the first System One model. Its Bkper model ID is jev.
- Versions.
jev-latestand versioned IDs are accepted but always select the current Jev; you cannot pin a version. Log the returnedmodel, and re-check tuned thresholds when it changes. - Input. Text only. English is Jev’s primary training language; other languages work with lower accuracy.
- Known weak spots. Literal reading, arithmetic, date comparison, and large states full of irrelevant detail. See Jev 1.13 jaggedness.
Differences from the TypeSafe API
Request and answer shapes match TypeSafe’s POST /v1/systemone, so TypeSafe’s guidance on questions, state, confidence, and patterns applies unchanged. Only the surroundings differ:
- Endpoint and authentication. Call
https://ai.bkper.app/v1/evaluationswith a Bkper access token. TypeSafe API keys and SDKs target TypeSafe’s endpoint and do not work here; use plain HTTP. - Errors. Bkper returns
400where TypeSafe returns422, and503withprovider_overloadedwhere TypeSafe returns529. A429can also mean your Bkper AI allowance is exhausted. - Usage. Requests count against your Bkper AI allowance, not a TypeSafe account.
Build with a coding agent
TypeSafe’s agent skill gives coding agents context on questions and patterns. When you use it for a Bkper integration, point the agent to this page — https://bkper.com/docs/ai/evaluations.md — for the endpoint, authentication, model ID, and errors.
Learn more
createEvaluationAPI reference — the field-by-field contract.- Models and Usage — evaluation models, limits, and usage rates.
- Bkper AI Gateway — access, tokens, and privacy.
- Add Bkper AI to an app — call evaluations from a Bkper Platform app.
- TypeSafe: System One, AI primer, Primitives, State, Confidence, and Patterns.