NVIDIA's Nemotron 3.5 Lightning Is Fast and Cheap, But It Ranks Near the Bottom for Accounting

Jan Abaza is the Founding PMM at DualEntry, where she shapes how the product is positioned for finance leaders moving off legacy ERPs. Her background spans product, design, and go-to-market, which gives her a practical read on what mid-market finance teams actually need. She holds an MBA from the University of Illinois Urbana-Champaign.

NVIDIA released Nemotron 3.5 Lightning on August 11, and the reaction was loud. It's a 30-billion-parameter open model that runs at a fraction of frontier-model cost, puts out tokens up to four times faster than rivals its size, and was built specifically to run the long-running agents that every enterprise AI team is trying to ship right now. Artificial Analysis measured it at 295 tokens per second and ranked it 14th out of 134 models on general intelligence, above the median for its size.
So we put it through our accounting benchmark. It scored 45.5%.
DualEntry's 2026 Accounting AI Benchmark runs 97 task-based questions across transaction classification, journal entries, accounts payable and receivable, bank reconciliation, financial reporting, and month-end close. Out of more than 40 models, Nemotron 3.5 Lightning came in second from the bottom. The top model, Grok 4.5, scored 84.2%. Several older and smaller models beat Lightning outright.
That split, strong on the general benchmarks and weak on the ledger, is the part worth paying attention to if you run a finance team.
What NVIDIA actually built
Lightning is a Mixture-of-Experts model. It has 30 billion parameters in total, but only around 3 billion of them fire on any given token, which is where the speed and the low price come from (roughly $0.05 per million input tokens and $0.20 per million output, well under frontier rates). NVIDIA is clear about the intended job: high-volume execution inside agent systems, meaning tool calls, result checking, and handing work off to subagents. The model even shipped with a router called NeMo Switchyard, and that router's whole purpose is to send routine execution to Lightning and push the harder reasoning up to bigger models.
So NVIDIA built Lightning to be the fast hands of an agent, not the brain. That is also why it does poorly on accounting.
Why speed is the wrong thing to measure for the ledger
Accounting doesn't reward "close enough, but quick." It rewards getting it exactly right. A single misclassified transaction, or a reconciliation that lands a cent off, is a failure, not a rounding error. The general benchmarks that make Lightning look good are measuring fluent reasoning and throughput. Our benchmark measures whether the model can post the right debits and credits inside a real chart of accounts, and it's graded by machine, with no partial credit for an answer that sounds confident but is wrong.
You can see the same pattern in the model's own numbers. Artificial Analysis noted that Lightning is unusually wordy, producing more than double the output tokens of comparable models during testing. When an agent is talking through its steps, that's useful. In a journal entry, it's just clutter. The things that make a model good at agent work are not the things that make it a careful bookkeeper.
Why it matters now
This is where teams get burned. Lightning is fast, open, and cheap, so it's an easy model to drop straight into a finance workflow. Someone wires it into an AP process or a close assistant because it costs almost nothing and answers quickly. But a model that gets fewer than half of real accounting questions right doesn't earn trust by being affordable. In finance, a wrong answer costs far more than the tokens that produced it.
That doesn't make Lightning a bad model. For the job NVIDIA designed it for, running the mechanical, high-frequency work inside agents, it looks strong, and the open license and low pricing will win it plenty of users. The mistake would be assuming that "good at agent execution" also means "safe for your books." It doesn't, and the benchmark shows how wide that gap really is.
This is why we test every major model release on real accounting work instead of trivia questions. The launch-day coverage will tell you which model is quickest and cheapest. It won't tell you which one you can trust to close your month.
See where Nemotron 3.5 Lightning and 40+ other models rank on the DualEntry Accounting AI Benchmark.



