Powered by Smartsupp

Probably Raises $9M From Andreessen Horowitz to Build Rigorous AI Hallucination Detection System



By admin | Jun 16, 2026 | 2 min read


Probably Raises $9M From Andreessen Horowitz to Build Rigorous AI Hallucination Detection System

Even the most advanced large language models continue to struggle with a persistent problem: hallucinations. These factual errors crop up in even the smartest systems, and while methods exist to detect them, the industry is still searching for the best solution. A startup called Probably, which recently secured $9 million in seed funding from Andreessen Horowitz, aims to establish a more rigorous approach to catching these mistakes. Founder Peter Elias (pictured above) explains that the company's mission is to stop hallucinations and simple factual inaccuracies from ever reaching the end user, targeting the kind of 99.99% accuracy typically seen in deterministic systems—a level far more elusive in AI.

Achieving such precision with LLMs, it turns out, demands a fundamental rethinking of many core AI engineering principles. Probably's first product is a data science tool designed to generate quick answers from complex datasets. Each result includes a citation and an audit trail detailing how it was produced—a practice that is becoming increasingly common among AI tools. However, preventing errors from slipping into these summaries required an elaborate harness system, which Elias describes as a "data science mech suit." The LLM's initial responses are checked against a deterministic validator system, which rejects any results that do not align with the dataset. Crucially, the LLM has been trained to work with this validator, and the entire system is optimized for both speed and accuracy, according to the company.

"What we learned building this was that the better your harness engineering is, the weaker the model can be," Elias says. "If you can refine the context enough, the model does not have to work very hard to do the right thing. Basically, it’s an exercise in reducing ambiguity."

This approach allows Probably's data science tool to run on significantly smaller AI models. Elias notes that the current version operates on a model "four classes weaker than the frontier models," meaning it can function on local hardware—a desktop computer rather than a data center—which dramatically cuts down on token costs associated with AI usage. This is a welcome development at a time when token expenses are rising and many companies are reevaluating their AI budgets. Elias's vision extends beyond data science; the same engine can be adapted for use cases like accounting or medical services—essentially, "any precision-sensitive use case."

"I think it’s really interesting that the big AI labs have not even attempted to do this," Elias says. "They’re incentivized not to, because they make money the more times you have to correct the model."




RELATED AI TOOLS CATEGORIES AND TAGS

Categories: Text Generation Art

Comments

Please log in to leave a comment.

No comments yet. Be the first to comment!