AI Fundamentals for Accounting
How language models and decision models behave, how to work with them, and how to keep numbers and decisions correct when accounting must be exact.
Most people meet AI through a chat box, and it feels like a calculator that talks. It is not. Accounting needs numbers that are correct, not approximately correct, and that takes a different mental model.
LLM Nature
A Large Language Model (LLM) does not look up answers. It predicts the most likely next word, piece by piece, with randomness baked in. Different models trade speed for depth, but all share this nature.
So the same prompt can produce different answers. Ask an LLM to compute a tax three times and you may get three slightly different numbers, and a fourth try can be off by more.
LLMs are probabilistic, not deterministic. Treat their direct output as a draft, never as a verdict.
Hallucinations
When an LLM does not know something, it does not stop. It guesses, fluently and confidently. It invents account names that don’t exist, tax rules that were never written, and invoices it has never seen. This is hallucination.
It is not a bug to be patched away. Chat models are trained to give answers people prefer, and that rewards confident-sounding guesses. The model’s uncertainty stays hidden inside fluent prose.
Never trust raw LLM output for facts. Verify it, ground it in your records, or route the work through something deterministic.
Context
Hallucinations get worse when the model has nothing to ground itself on.
An LLM only knows what is in its context window right now: your prompt, the files you attach, and the recent conversation. It does not know your Books or remember last week. Each session starts blank.
You build context by handing it relevant pieces, such as a chart of accounts, a transaction list, a policy document, or a project’s AGENTS.md. Persistent context, like skills and project files, saves you from pasting the same information every time.
Context has a sweet spot. Too little, and the model invents. Too much, and it loses focus and mixes unrelated pieces. Curate: give the model what it needs for the question in front of it, nothing more.
The curve never touches zero: better context lowers hallucination risk, but it does not make the model deterministic.
Intent
Intent means telling AI what done looks like, not listing every step. A useful request names four things:
- Outcome — what you need.
- Reason — what it is for.
- Source of truth — where the facts come from.
- Success criteria — how you will both know it is right.
Step-by-step: “Open my book, filter transactions tagged
#salesfor Jan–Mar, sum the VAT column, convert to EUR, give me the total.”
This scripts the work. The model can still misread a step, skip one, or invent around it.
Intent: “I need the VAT I owe for Q1 2025, in EUR, ready to file. Use my Bkper book as the source of truth.”
This describes the outcome. The agent decides which transactions to pull, which math to run, and whether to answer directly or write a small script.
Pair intent with a concrete check: an expected total or range, a report that matches last quarter’s shape, a reconciliation that should come out to zero, or an Account whose closing balance you know. Without one, the model cannot know when it is done, and neither can you.
Agents
An agent is an LLM running in a loop with tools. At each step the model chooses an action, runs a tool such as a CLI command, a script, or an API call, and observes the result. The observation shapes the next step: progress, a correction, a retry, or a different approach. The loop ends when the success criteria are met.
This is the shape behind Bkper CLI Agent, coding agents, and other tool-using assistants. Success criteria close the loop; without them, a probabilistic engine produces drift, not progress. And a loop is only as trustworthy as the tools inside it.
Agents work best when a person is watching: building a script, exploring a question, preparing work for review. But the model chooses every step, and each step is another chance to go off course. Work that repeats thousands of times without supervision needs a different shape.
Decision Models
Not every model writes text. A decision model doesn’t chat; your code calls it. It reads the facts you hand it, such as a bank line and your list of Accounts, and answers in a format you set in advance:
- Is this true? → a probability of yes, from 0 to 1.
- Which one of these? → one of the options you listed, with a confidence score.
- Where on this scale? → a position on levels you described, with a confidence score.
It is still probabilistic, but it is trained for calibrated decisions: it reports its uncertainty as a number, instead of hiding it in fluent prose.
That makes a different shape possible: code owns the workflow, and the model answers narrow questions at decision points. The model evaluates; your code decides.
When a bank line arrives, code asks which Account should receive the money. If confidence is above your threshold, code records a normal transaction from Checking to that Account, so the Book stays balanced whatever the model answers. Otherwise, a person reviews it.
Decision models don’t calculate: amounts, dates, and balances stay in code. Bkper’s first decision model is Jev, by TypeSafe. Learn more in Decision Models.
AI in Accounting
Accounting cannot be 99% right. A balance sheet that is mostly correct is wrong, and a tax filing that is approximately accurate is a problem. And no technique — better prompts, richer context, smarter agents — makes an LLM’s output guaranteed correct. Errors will happen, and inside an agent loop they compound silently between checks.
So the rule is not make the AI correct. Nothing makes the AI correct. The rule is:
Never let unverified LLM output be the final word on a number.
Code keeps verification cheap. When an LLM writes a script that computes the answer, you stop verifying outputs and start verifying the script. Read it once, test it, and trust it as long as it doesn’t change. The same inputs then give the same outputs, auditable line by line. Verification becomes a one-time cost instead of a per-result cost.
Code carries the calculations, and decision models carry the judgments inside that code. Each kind of accounting work has its route:
| Work | Route | A person reviews |
|---|---|---|
| Calculations: taxes, balances, reports, reconciliations, financial statements | An LLM writes code, or you use a deterministic tool | The code, once |
| Operational decisions: categorizing bank lines, matching duplicates, flagging entries | Code asks a decision model and acts only above thresholds you set | The code and thresholds once, then the uncertain cases |
| Open-ended work: business insights, a first chart of accounts, period summaries | An LLM drafts | Every draft |
The routes combine. A coding agent can write the whole workflow: code for the math, decision-model calls for the judgments, and thresholds for when to ask a person. The LLM builds the system once instead of sitting in the loop every time it runs.
AI doesn’t remove the reviewer; it changes what arrives for review. A person reads the code once and handles the cases it flags as uncertain, instead of re-checking every number and decision a model emits.
Further watching
- “Never Trust An LLM” by Matt Pocock — a developer-oriented explanation of why LLM output must be verified instead of trusted directly.
What’s next
- Use Bkper with your AI assistant — connect the assistant you already use and get a first useful result.
- Bkper CLI Agent — work in a terminal agent with Bkper context built in.
- Coding Agents — build Bkper integrations with grounded coding agents.
- Decision Models — ask a decision model bounded questions and turn the answers into decisions in code.
- Docs for AI — get Bkper docs and context into AI tools.