Most AI underwriting products are built the same way: a general-purpose model, a generic orchestration layer, and the assumption that this is enough.
By
Adam Ben-David
·

By Adam Ben-David, Senior Director of AI, hyperexponential
Most AI underwriting products are built the same way: a general-purpose model, a generic orchestration layer, and the assumption that this is enough.
In insurance pricing and underwriting, it isn't.
McKinsey estimates generative AI could unlock $50–70 billion in additional annual revenue for the insurance industry. By 2030, more than 90% of pricing and underwriting for many policies could be automated. But the value comes from rewiring how decisions get made, not from swapping in a smarter model. McKinsey's own analysis puts 60–80% of AI value in insurance in "traditional" AI (predictive pricing, risk modeling, automation) and only 20–40% in generative AI. The emphasis is on data, workflows, and governance. Not model selection.
If the model isn't the differentiator, what is?
The parts that matter (and the parts that don't)
An agent system has four components: the LLM, the agent loop (how it iterates toward a goal), context (domain-specific information), and tooling (what actions it can take).
Two of these are commodities now. Frontier model performance has converged across providers and price points. Agent orchestration frameworks have settled on common patterns for task management, memory, and error handling.
The value is elsewhere. McKinsey estimates AI technologies could add up to $1.1 trillion in annual value for global insurance, with roughly $400 billion from pricing and underwriting. Customer service and personalization could contribute another $300 billion. All of that is framed as decision and workflow improvement. None of it depends on which foundation model you pick.
Generic models operating without domain grounding actively mislead. An empirical study of LLMs in finance found that off-the-shelf models exhibit "serious hallucination behaviours" on financial terminology and numerical tasks, even when exact answers exist.
That leaves context and tooling.
What a generic agent doesn't know
Take a commercial property submission. The underwriter needs to apply appetite guidelines against current portfolio position, interpret loss history for the specific risk, and form a view on technical price given the market, broker dynamics, and renewal timing.
The FAITH benchmark, which tests LLMs on real financial tables, found that even top-tier models show 10–20% error rates on multi-step numerical reasoning. In underwriting, a 10% error rate on pricing arithmetic is a portfolio problem.
hx's Ingestion Agent turns messy submission documents into structured, usable data. It reduces ingestion time from hours to as little as five minutes on complex lines of business by mapping directly into the hx model schema.
The Actuarial Agent runs inside the hx IDE. It understands the framework and source code, and can generate data schemas, UI components, and rating logic from natural-language prompts and existing Excel raters. It accelerates model build and refinement by up to 10x. That matters because 99% of actuaries say their technology still needs improvement. 85% still rely on Excel, and 48% cite lack of version control as a key blocker. The desire for technology that can create more accurate models faster rose from 39% to 70% year over year.
The Underwriting Agent connects submission analysis to quote generation while natively understanding your pricing models. It uses your actual logic directly, not an approximation of it.
These agents work because they're connected to the context that matters: your models, your rules, your portfolio.
Where vertical agents pull ahead
Vertical advantage comes from two places.
Domain context. An agent's usefulness scales with how well it understands your underwriting appetite, pricing philosophy, and portfolio constraints. You can't buy this from a horizontal platform. It has to be built by people who understand how underwriting decisions are made.
Insurance-native tooling. Connecting an agent to the systems where underwriting data lives requires domain knowledge at the integration layer, not just the model layer. Each connection has to be built and maintained correctly.
These compound. An agent with better context uses its tools more accurately. More accurate tooling produces outputs the agent can reason over without correction. Every decision the agent supports adds to the institutional knowledge it draws from next time.
Closing this loop requires a live feedback mechanism where teams can see how pricing and underwriting decisions are shaping portfolio performance in near real time. Portfolio steering becomes continuous rather than a once-a-year exercise. Instead of infrequent batch rating, teams can test changes more often, run experiments, and act on what the data shows. Those performance signals feed back into the governed context agents rely on, so decision quality improves with use.
McKinsey's analysis of AI in insurance investing supports this pattern. Investors that systematically prioritize operational value creation, including AI-enabled decisioning, achieve internal rates of return 2–3 percentage points higher than peers. The compounding effect is real and measurable.
The right question when evaluating agents
Most insurers evaluating agents focus on which platform to use. That's the wrong starting point.
The underlying model will keep improving. That part takes care of itself.
The FINOS AI Governance Framework reinforces the architectural point: decision quality depends on how well agents are grounded in governed data sources, and how reliably they can retrieve and act through well-designed tooling. McKinsey's investor analysis makes a similar point: insurers are moving to modular, interoperable architectures, and vendors who "facilitate connectivity between data, models and automated agents" are likely to become key infrastructure.
The real question: what context and tooling will your agents have access to, and who built it? That's where decision quality comes from. Generic agents are interchangeable. The teams that encoded your underwriting logic, and the depth at which it's wired into your systems, are not. That gap keeps widening.

Adam Ben-David



