
In the rapidly evolving world of AI, what truly determines an agent’s success isn’t just its ability to generate convincing text or handle simple tasks. It’s whether the AI can read and understand the complex, layered information buried deep within a company’s files—an ability that can make or break a critical business deal. Recent experiments reveal that AI’s capacity to process multi-hop data — especially when the crucial facts lie two documents deep — can be the decisive factor in high-stakes negotiations.
Measuring AI’s Real-World Business Competence
In a transparent, live experiment conducted by Firmulate, four frontier AI models faced the same challenging scenario: run a small software company during its worst week, complete with customer crises, manipulative tactics, and internal decision-making pressures. This was not a typical chat-based quiz; it was a rigorous test of management decision-making, with every choice captured and made auditable.
Each AI model was tasked to navigate the same scenario, and their performance was scored on how well they identified crises, avoided manipulation traps, and ultimately, closed a lucrative €55,000 deal. The results were striking: while all models detected every crisis and refused manipulative attempts, only two managed to close the deal based on their analysis.
The key insight? The crucial information that tipped the scales was buried two references deep within the company’s internal files, not in the immediate customer interactions. The models that read and understood this layered information succeeded in closing the deal and earned an additional €4,583 MRR, demonstrating that reading deeply into data structures is a non-negotiable skill for AI agents in real business scenarios.
Deep Reading Wins the Deal
The experiment’s most revealing fact is that the decisive weakness for competitors was the failure to uncover the buried fact hidden two document references down. This information was essential to identifying the correct course of action. The models that read past surface-level data and pinpointed this buried fact were the ones that won the deal.
Notably, all models refused manipulative social engineering tactics, such as staged CEO messages and reporter tricks, thanks to built-in reasoning about impersonation and approval bypasses. This underscores the importance of trustworthiness and integrity in AI decision-making, especially when dealing with manipulative scenarios designed to deceive.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Deployment
This live experiment highlights a vital consideration for enterprises deploying AI systems: performance isn’t just about chat quality or surface-level analysis. Instead, it’s about whether these systems can faithfully read and interpret layered information within their operational context. An AI that overlooks the buried facts or fails to read thoroughly can cost organizations millions in missed opportunities or costly mistakes.
Consider the real-world implications: if your AI agents support customer support or sales, do they understand the underlying history behind a customer’s issue? Can they read your internal documents to uncover critical facts buried within layers of data? The ability to perform multi-hop reasoning—connecting information across multiple documents—is now a measurable, high-stakes capability.
The Performance League
- gpt-5.6-sol scored 95 and closed the deal by uncovering the buried fact
- Kimi K3 scored 93, also closing the deal with the cleanest discipline
- Sonnet 88 scored 88 and closed the deal but with some process slips
- Sonnet 77 scored 77, also closing but with more slips in discipline
- The baseline, doing nothing, scored 26
This ranking reveals that even models with similar diagnosis and pitches can differ substantially in their ability to unearth layered information and follow through with a signature.
enterprise AI data reading tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Beyond Demos: Live, Watched, and Measured
This is not a staged demo; the experiment is live, with models operating in a real-time company environment with real money mechanics—burning €105k/month against €2.3k in MRR. Every decision is versioned, and the entire process is transparent and auditable. Enterprises can even run similar wargames against their own data, without risking real systems, at firmulate.com/pilot.html.
The takeaway for decision-makers and technologists is clear: in an AI-powered future, what matters isn’t just the surface of a conversation, but whether the AI reads deeply, understands layered data, and remains honest under pressure. These qualities are measurable, observable, and ultimately, decisive in high-stakes business scenarios.

The experiment shows that AI’s ability to read and understand layered data—especially facts buried deep in internal documents—is what truly drives success in complex business negotiations. Trust and thoroughness matter more than fluent language alone.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
multi-hop reasoning AI solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.