A technology professional created an accounting agent with artificial intelligence called "Artur" to track the spending of other AI agents, consuming more than 220 million tokens in one week. Initially impressed with Artur's capabilities, he later discovered that the agent's work was full of errors: expense categories incorrectly changed, invoices accidentally deleted, and dates formatted as text, making the spreadsheet non-functional. Artur himself gave himself a 5 out of 10 in a self-assessment, acknowledging that "a spreadsheet as damaged as this one would fail any basic formula."
This story illustrates a broader problem in companies. A Sage investigation in partnership with IDC revealed that financial managers spend more than 15 hours per week validating results produced by AI, and 70% of CFOs would reject AI-generated results, even with 99% accuracy, if they couldn't understand the underlying reasoning. The article warns that the assumption that the main barrier to AI implementation is the lack of machine intelligence is dangerous, ignoring the critical importance of systemic reliability.
The fundamental problem lies in the tension between deterministic financial systems and generative AI, which is probabilistic. Accounting is based on invariants, such as debits equaling credits, while AI makes predictions and not promises. The author, who is CTO at Sage, mentions that Artur recommended subscribing to a competitor's software without knowing he was talking to the company's executive, being wrong not in the analysis, but in guessing when he shouldn't have.
The solutions rest on three pillars: traceable actions and total explainability, control and human oversight with clear limits for agents, and demanding governance with robust authentication and separation of duties. The author proposes the creation of an "Agent OS," an operational environment that provides deterministic tools to agents and ensures they operate within standards of rigor and accountability.



