Elon Musk's Grok Beat ChatGPT & Claude at Accounting (2026)

Photo of Jan Abaza, DualEntry blog author
Jan Abaza
Founding Product Marketing Manager, DualEntry
Photo of Jan Abaza, DualEntry blog author
Jan Abaza
Founding Product Marketing Manager, DualEntry

Jan Abaza is the Founding PMM at DualEntry, where she shapes how the product is positioned for finance leaders moving off legacy ERPs. Her background spans product, design, and go-to-market, which gives her a practical read on what mid-market finance teams actually need. She holds an MBA from the University of Illinois Urbana-Champaign.

Learn about our editorial policies.
Last updated
August 3, 2026
Reviewed by
Do San (Justin) Myung
Do San (Justin) Myung
Do San (Justin) Myung
Expert Accountant & Former Consulting CFO | DualEntry

Justin (Do San Myung) is Expert Accountant at DualEntry with 20+ years of hands-on experience managing general ledgers, financial close processes, and ERP implementations for mid-market and enterprise companies. As a former Consulting CFO and Controller, he has personally overseen month-end closes, SOX compliance programs, and multi-entity consolidations across technology, manufacturing, and services industries. Justin specializes in transforming manual accounting workflows into automated, AI-driven processes.

Learn about our editorial policies.
2026 accounting AI benchmark chart: Grok 4.5 leads 42 models at 84.2% accuracy, with no model clearing 85%.
Contents
More

Subscribe to the
DualEntry Newsletter

Get Fresh Al finance insights, reports and more delivered straight to your inbox

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Summarize this article

Independent tests rank Grok 4.5 around fourth on general intelligence. On real accounting, it beat every one of the 42 models we ran.

Grok is the AI people argue about. It makes headlines for the things it says, not for its way with a spreadsheet. So it was odd to watch it quietly win a test that rewards the opposite of all that: simply getting the numbers right.

We ran 42 of the leading AI models through a battery of real accounting work, not multiple-choice trivia but the tasks a staff accountant actually clears in a day. Grok 4.5 came out on top at 84.2% accuracy, ahead of models from every major lab. Second place went to something almost as unexpected: an open-source model.

The timing sharpens the point. Grok 4.6 is landing this week with no benchmarks attached - SpaceXAI still hasn't published scores for it - which leaves independent tests as the only real read on how good these models are right now. Ours is one of them, and it is the only one built entirely on accounting.

The top of the leaderboard

Here is where the leaders landed on overall accuracy.

Rank Model Provider Accuracy License
1 Grok 4.5 SpaceXAI 84.2% Closed
2 Z.ai GLM-5.2 Zhipu AI 83.2% Open
3 Claude Fable 5 Anthropic 83.2% Closed
4 Gemini 3.5 Flash Google 81.7% Closed
5 GPT-5.5 Pro OpenAI 80.2% Closed
5 Gemini 3.6 Flash Google 80.2% Closed
5 Claude Opus 4.6 Anthropic 80.2% Closed

One point separated first from third. Grok won, but it won a close race, not a runaway.

Why this is a real surprise

Grok's reputation has almost nothing to do with careful, rules-bound work. It is the model built to be provocative. Accounting is about as far from provocative as work gets. Debits have to equal credits. Accruals reverse on schedule. A reconciliation ties out or it doesn't, and no amount of personality moves the number.

It's a surprise in the rankings, too. When the independent firm Artificial Analysis scored Grok 4.5 on general intelligence, it landed around fourth, behind Claude Fable 5, GPT-5.5, and Claude Opus 4.8. Put the same model in front of real accounting work and it finishes first. General-purpose smarts and domain reliability are clearly not the same measurement.

So it's worth sitting with the fact that a model famous for stirring things up turned out to be the steadiest hand at double-entry bookkeeping. What an AI is known for and what it's quietly good at can be two completely different things.

The second surprise: open source is right behind

The runner-up, Z.ai's GLM-5.2, is an open-weight model, and it tied Anthropic's closed flagship, Claude Fable 5, at 83.2%. A handful of other open models, including GLM-5.1, DeepSeek V4 Pro, and Kimi K3, placed inside the top 12.

For a finance team, that is the more useful story. Open models can run inside your own environment, under your own controls, without shipping a general ledger off to someone else's servers. Now that they sit within a point of the best closed systems, the build-versus-buy question for sensitive financial data looks different than it did a year ago.

The number underneath the winner

Look past the trophy and the figure that matters is this: nobody cleared 85%.

The best accounting AI on the planet still misses roughly one task in six. In accounting, that missed sixth is rarely harmless. It shows up as a misclassified transaction, an entry that won't balance, or a reconciliation that drifts a little further from reality each month until someone catches it at close.

Grok won, and that's a real result. But no CFO should read it as permission to hand over the books. Read it as confirmation that these models have become useful assistants to a person who still owns the review. They are good enough to handle a first pass and nowhere near good enough to let you skip the second.

A higher version number is a release date, not evidence

One more note for anyone who assumes the newest model is automatically the best one. Even within a single lab, higher version numbers didn't track with higher scores. Claude Opus 4.6 scored 80.2% and beat both Opus 4.8 (75.3%) and Opus 5 (72.3%).

How we tested

We gave each model a provisioned chart of accounts and 101 task-based questions across eight categories built to mirror a real workflow: transaction classification, journal entry creation, accounts payable, accounts receivable, bank reconciliation, financial reporting, month-end close, and applied accounting knowledge.

Grading is deterministic. A task either produces the correct result or it doesn't, with no partial credit for sounding confident. Every model runs in an isolated environment with no connection to a live account, and each benchmark runs several times so we can report accuracy, standard deviation, and a difficulty tier for every category.

The full leaderboard is interactive. You can filter it by provider, license type, and more, and check our math yourself. See the full 2026 Accounting AI Benchmark →

The bottom line

For now, the most talked-about AI on the internet is also the most accurate at accounting. It still gets one task in six wrong. That is where AI in finance stands in 2026: a strong co-pilot that still needs a pilot in the seat next to it.

Finance leaders already sense this. In recent surveys only 14% of CFOs say they completely trust AI to produce accurate accounting data on its own, and 97% still call human oversight critical — part of why finance is still among the slowest functions to hand work to AI. A benchmark that tops out below 85% is exactly why that caution is earned.

DualEntry Labs runs this benchmark on real accounting workflows and updates it as new models ship. Want the next round in your inbox? Subscribe here.

Sources

  1. Grok 4.6 launch and benchmark status: kie.ai
  2. Independent intelligence ranking of Grok 4.5: Let's Data Science / Artificial Analysis
  3. CFO trust and oversight data: GlobeNewswire, Journal of Accountancy
  4. All model scores from DualEntry's live benchmark leaderboard.

See the full power of DualEntry in 30 minutes

Go live in 24 hours

By clicking "Schedule Demo" you agree to the use of your data in accordance with DualEntry's Privacy Notice, including for marketing purposes.