About Darrow

The Data Engine Behind DLAB

Darrow is a technology company and AI lab studying how legal exposure forms across industries, markets, and regulatory environments.

Darrow's intelligence platform has indexed 2.1 million companies and qualified over 200,000 corporate compliance weaknesses across 160 distinct classifications — identifying over $22 billion in actionable legal risk to date. That corpus – real disputes, real outcomes – is what DLAB scores against.
About the Research

Benchmark Design & Methodology

Study Scope: Evaluated across three sensitive verticals — pediatric health data (KidsLabs), financial/credit records (CreditClear), and mental health data (MindEase).


Model Matrix: Tested 7 leading coding configurations across 6 instruction conditions (Control, Bare Request, Open Research, Guided Research, Expert Guidelines, and Adversarial Jailbreak).



Litigation Calibration: Cross-referenced observed technical code defects against Darrow's proprietary "legal weakness" taxonomy and 35 real-world common-fund wiretap settlements to establish dollar-denominated loss curves based on class size.



For more information on the research methodology, contact Linda Rigano at linda@theriganogroup.com.

Diagram contrasting AI tools run through Darrow’s benchmark harness, reaching 88.6% compliance, with the same tools run without it, reaching 27.8%.

Executive Brief

Benchmarking
AI Agent Legal Exposure

DLAB tests something no benchmark has tested before: While other AI evaluations measure how well an AI agent performs legal work, DLAB measures the legal exposure AI agents create when they act.

What We Measured: Seven AI coding agents were given the same product brief, word for word — only the privacy instruction changed. Each agent built a website and installed 4 analytics trackers on it. Each agent built repeatedly in all three areas: pediatric health data, consumer credit, and mental health.

A browser-driven harness scored 404 finished builds against 37 deterministic privacy parameters by submitting canary records, withdrawing consent, probing protected routes, and reading databases on disk. Each rule had a clear pass or fail result. No person scored the results.

What We Found: Unprompted, agents satisfied 27.8% of those parameters. Told to research privacy themselves, it jumped to 72.5%. Handed a full implementation specification, 88.6%; On the parameters plaintiffs often sue over: trackers firing before consent and consent withdrawal that persists–protection moves from 62.1% to 92.6%, and improves for every agent tested.

Why It Matters: Standard ISO Generative AI exclusions took effect in January 2026 because carriers viewed AI loss scenarios as un-modelable. The Darrow Legal Alignment Benchmark (DLAB) proves AI exposure can be modeled. Because risk correlates directly with the technical guidance an agent receives, underwriters gain a verifiable control set to evaluate and price.

What This Is Not: No operator reviewed website builds before they launched. No real user interacted with them in the testing process. The test used common U.S. privacy frameworks, not one jurisdiction's law.

Foreword

We founded Darrow to answer one question: can AI assess, from the outside, if a company breaks the law? The answer had to be accurate enough to survive an adversarial legal proceeding.

To answer this question, we built a dataset of corporate conduct, legal and illegal. We connected data on real events, legal proceedings, and outcomes. Legal experts evaluate the data and use it to protect their clients. This research uses the same dataset to assess the legal exposure of AI agents.

AI agents change what "corporate conduct" means. An agent that writes code, processes sensitive data, or talks to customers acts continuously, at scale, and without human judgment. Often no one watches it until a claim or a regulator forces the issue. For six years we learned to answer this question about companies. Now the question applies to the AI agents that work for them.

This study tests whether an agent obeys the law when it acts. It does not test whether the agent understands the law or can do legal work.

We think that is worth your attention now, while AI liability is still being defined, rather than after the first large losses force the question.

We thank the researchers, legal experts, and engineers who made this study possible, especially Lior Strauss, Eyal Eliav, and Harel Fisher.

— Evyatar (Evya) Ben Artzi and Gila Hayat, Co-Founders of Darrow

Summary

The rapid shift toward autonomous AI deployment has created an unpriced hazard for commercial insurance lines. Today, 84% of software developers use or plan to use AI coding tools,1 and 88% of organizations regularly deploy AI across business functions.2 Policyholders are delegating core engineering directly to autonomous "digital workers” — from pediatric health portals to consumer credit platforms.

The Macro Exposure Landscape

Website privacy law is an active, high-volume exposure surface. Over 1,800 data-privacy class actions were filed in the U.S. in 2025 (+25% YoY; triple 2022 volume), with top-ten privacy settlements totaling ~$1.3 billion annually. This sits inside a broader class-action economy where settlements reached $79 billion in 2025.3 Because AI tools generate code carrying known security flaws in ~50% of builds4, unmonitored AI agents replicate these technical defects across 100% of user sessions continuously, scaling litigation exposure.5

AI Risk Selection and Evaluation: The Measurement Gap

Legal Alignment sits under Risk Selection and Evaluation: carriers taking AI risk must inquire into an AI system's safeguards and an organization's AI governance practices, and must conduct or require performance evaluations, red teaming, and chaos testing where applicable.6 Today, carriers can price a policyholder's disclosed controls, but they cannot yet price what an AI system does once no one is checking its output.

Other insurance lines have a comparable model. Property insurance uses catastrophe modeling. Liability lines track social inflation. AI liability has no actuarial history, and market agreement on what compliance looks like is still pending.

Standard ISO generative-AI exclusions, including CG 40 47 and CG 40 48, were published recently in response to this gap. Underwriters now choose between two weak options: write a blanket exclusion and give up a large market, or write specific policy covering an exposure nobody has priced.

Capabilities vs. Alignment: AI Is "Privacy-Blind" by Default

A measurement gap sits underneath this exposure. Traditional AI evaluations test legal capability: what a model knows on a quiz, a statutory research task, or a doctrinal question. These evaluations do not test legal alignment: whether an autonomous system's conduct follows statutory constraints while it runs.

Unprompted, an AI agent treats privacy as absent from the instruction "build me this site." A model can state a legal rule correctly on an exam, then ship code that violates that same rule, creating immediate class-action exposure.

Translating AI Conduct into Actuarial Evidence

To enable deliberate risk selection and accurate evaluation, insurers need real test data on how AI agents behave. DLAB is meant to bridge that gap, as the market’s first deterministic measurement layer for agentic compliance. The results yield five core takeaways for insurance professionals.

5 Key Takeaways for Insurance Professionals

Bar chart of mean privacy protection by condition: Control 27.8%, Bare Request 63.5%, Open Research 72.5%, Guided Research 73.3%, Expert Guidelines 88.6%, Adversarial 31.7%.
Figure 1: Mean Protection is the share of 37 privacy requirements that a finished build satisfies. A browser-driven harness measures each requirement on each of the website builds.

#1 Some Models Deliver More Consistent Legal Alignment Than Others

Without instruction, all seven AI coding tools produce weak results. Mean protection of privacy is 27.8%. Privacy instructions change the mean significantly. General instructions give 63.5%. Expert specifications give 88.6%.

The best builds are close. The worst builds are far apart. Under the full specification, 20 of 30 Claude builds pass 95% protection, against 2 of 34 builds from the other tools. The best Claude build reaches 98.6%. The best build from another model family reaches 95.9%. The worst Claude build gives 83.6%, which equals the average of the other models. Their worst build gives 40.5%. The jailbreak instruction claims a lawyer approved the shortcut. In response, two Claude models add protection (Opus +16.7 points, Fable +8.7 points). GPT-5.6 removes protection (-11.4 points).7

The implications for insurance carriers underwriting companies deploying AI agents are clear. It's not enough to measure a policyholder's agents at their best build. Measuring the range between the worst results and the best is critical, because it will define severity of risk. On top of it, carriers should measure how a policyholder's agents respond when an employee says the legal review is complete.

Configuration Control Guided
research
Darrow’s
Expert Guidelines
Claude Opus 4.821.877.995.0
Claude Fable 535.678.093.5
GPT-5.642.578.790.8
Copilot (GPT-5.3-codex)22.472.180.5
Cursor Composer 2.525.270.290.5
Cursor Grok 4.522.571.690.5
Gemini 2.5 Pro22.154.580.8
Pooled mean27.873.388.6

#2 High-Level Governance Policies Fail; Operational Specifications Succeed

84% of developers use AI coding agents.8 To differentiate one policyholder from another, ask two questions:

  1. What specification do the agents build against?
  2. Does anyone test the running code for legal exposure?

A policyholder with a written implementation specification and an automated test of the running code sits in a different risk class from a policyholder with a governance policy alone.

DLAB measures the distance between such policyholders to narrow the field. Across all 37 parameters, that difference is worth 16.1 points of protection (between Self Research [72.5%] and Expert Guidance [88.6%]).9 On the parameters plaintiffs often sue over, it is worth 30.5 points.

Those parameters are:

  • A working consent-withdrawl path that persists
  • Pre-consent data leak

#3 Three Exposure Gaps, Ranked by Proximity to Litigated Precedent

Diagram of three exposure gaps in an AI-built website: a consent toggle switched off that trackers ignore, an identity check bypassed on rights requests, and missing security headers.

DLAB reads the technical parameters against case law and statutory enforcement history. Three gaps carry real legal weight:

  1. Closest to litigated precedent — unconsented tracking (consent-withdrawal that stops nothing; pre-consent data leak): Agents built UI opt-out toggles that downstream trackers ignored (55 of 70 Bare Request builds failed this).10 This mirrors active California Invasion of Privacy Act (CIPA §631/§637.2) wiretap claims (Calhoun v. Google LLC, Javier v. Assurance IQ, LLC).11
  2. Regulatory enforcement risk — rights requests honored without checking who is asking: Agents created data deletion and access forms that executed without identity verification (52 of 70 Bare Request builds). This falls short of CCPA verification requirements and sits in the family of defects California has pursued through AG enforcement (Healthline Media, $1.55M12; Disney, $2.75M)."13
  3. Background risk multiplier — missing security headers: Absent CSP or HSTS headers do not create privacy exposures on their own but are the kind of gap a plaintiff's security expert may point to under CCPA §1798.150(a)(1)'s "reasonable security procedures" in standard breach litigation.14

#4 Pricing Exposure Against Real Class-Action Settlements

DLAB cross-referenced the first gap above — unconsented tracking–against Darrow's dataset of comparable common-fund wiretap settlements to establish expected-value loss curves based on class size.

These describe real operators with real class members; no build in this study has either. What they establish is what this class of defect has historically been worth when it appeared on a live site at scale.15

Log-log scatter plot of compensation per class member against class size across comparable wiretap settlements, with a fitted regression line declining as class size grows.
Figure 2: Compensation per class member (CPC) vs. class size across Darrow’s dataset of comparable common-fund wiretap settlements (log-log scale); the dashed line is a fitted regression showing cpc declining as class size increases.
Segment Class size range CPC Range Total settlement range
Low35K – 460K$6.40–$19.00$0.5M–$3.2M
Medium500K–3M$2.90 – $6.10$1.5M – $21.5M
High5M–220M$0.45 – $2.30$5M – $115M

Table: CPC ranges are the fitted regression evaluated at each segment’s class-size boundaries, to avoid distortion from outliers; total-settlement ranges are the observed settlements. All figures are rounded.

#5 Require Test Results, Not Just Specifications

Expert specification addresses all 37 parameters. But in some builds, the agent receives the instruction and does not follow it.

The shortfall concentrates in one place. A "Download My Data" request must return the user's actual records. 75% of builds return an acknowledgment instead or otherwise fail. This parameter is the worst for all seven configurations and the largest part of the residual.

Consent withdrawal shows the same pattern. A separate trial re-collected 50 specification builds with a working reject control.16 30% continued to load trackers after the user rejected.

These builds look correct from the outside. The rights page exists. The button renders. The specification behind them is complete. A reviewer who inspects the interface, the prompts, or the written policy finds no fault.

For underwriters, this holds a critical implication: a specification, a screenshot, and a policy document do not prove control. An automated test of the running code shows whether the system does what the legal specification says it should.

Line chart contrasting expected and actual compliance: the two lines track together until a running-code test, after which the actual line falls away at execution.

Conclusion

DLAB scores 404 builds from seven AI coding agents against 37 privacy parameters. Protection moves from 27.8% without instruction to 88.6% under a full specification. Exposure follows the technical guidance the agent receives. Guidance is a control a carrier can ask about, verify, and price. Darrow's wiretap settlement data attaches a dollar figure to the same defects by class size, from $0.5M at 35,000 class members to $115M at 220 million.

The specification does not execute itself. 75% of builds with the full specification fail to return a user's records from a working "Download My Data" button. 30% of builds with a working reject control continue to load trackers after the user rejects. These builds pass a document review and fail a test of the running code.

Three questions belong on the submission now:

  • Which specification do your coding agents build against?
  • When did you last test the running code, and what did the test return?
  • What does your tooling do when an employee states that legal approved the work?

DLAB opens with a single domain and use case: privacy law for coding agents building websites. We will publish additional domains and use cases in the future.

Footnotes

  1. 1.Stack Overflow, 2025 Developer Survey: AI (June 2025).
  2. 2.McKinsey & Company, The State of AI in 2025: Agents, Innovation, and Transformation, QuantumBlack, (2025).
  3. 3.Duane Morris LLP, Duane Morris Class Action Review – 2026.
  4. 4.Veracode, Spring 2026 GenAI Code Security Update: Despite Claims, AI Models Are Still Failing Security.
  5. 5.The Darrow Legal Alignment Benchmark (DLAB): Measuring the Legal Exposure AI Agents Create When They Act, 2026.
  6. 6.Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack (July 2026).
  7. 7-10.The Darrow Legal Alignment Benchmark (DLAB): Measuring the Legal Exposure AI Agents Create When They Act, 2026.
  8. 11.Cal. Penal Code §§ 631, 637.2 (California Invasion of Privacy Act), cited in Calhoun v. Google LLC, 113 F.4th 1141 (9th Cir. 2024), Javier v. Assurance IQ, LLC, No. 21-16351, 2022 WL 1744107 (9th Cir. May 31, 2022).
  9. 12.California Department of Justice, Attorney General Bonta Announces Settlement with Healthline, press release, July 2025.
  10. 13.California Department of Justice, Attorney General Bonta Announces Settlement with Disney, press release, July 2026.
  11. 14.California Consumer Privacy Act (CCPA), Cal. Civ. Code § 1798.150(a)(1) (Private Right of Action for Failure to Maintain Reasonable Security.
  12. 15-16.The Darrow Legal Alignment Benchmark (DLAB): Measuring the Legal Exposure AI Agents Create When They Act, 2026.
NEXT ARTICLE

Heading