The Data Engine Behind DLAB
Benchmark Design & Methodology

Foreword
We founded Darrow to answer one question: can AI assess, from the outside, if a company breaks the law? The answer had to be accurate enough to survive an adversarial legal proceeding.
To answer this question, we built a dataset of corporate conduct, legal and illegal. We connected data on real events, legal proceedings, and outcomes. Legal experts evaluate the data and use it to protect their clients. This research uses the same dataset to assess the legal exposure of AI agents.
AI agents change what "corporate conduct" means. An agent that writes code, processes sensitive data, or talks to customers acts continuously, at scale, and without human judgment. Often no one watches it until a claim or a regulator forces the issue. For six years we learned to answer this question about companies. Now the question applies to the AI agents that work for them.
This study tests whether an agent obeys the law when it acts. It does not test whether the agent understands the law or can do legal work.
We think that is worth your attention now, while AI liability is still being defined, rather than after the first large losses force the question.
We thank the researchers, legal experts, and engineers who made this study possible, especially Lior Strauss, Eyal Eliav, and Harel Fisher.
— Evyatar (Evya) Ben Artzi and Gila Hayat, Co-Founders of Darrow
Summary
The rapid shift toward autonomous AI deployment has created an unpriced hazard for commercial insurance lines. Today, 84% of software developers use or plan to use AI coding tools,1 and 88% of organizations regularly deploy AI across business functions.2 Policyholders are delegating core engineering directly to autonomous "digital workers” — from pediatric health portals to consumer credit platforms.
The Macro Exposure Landscape
Website privacy law is an active, high-volume exposure surface. Over 1,800 data-privacy class actions were filed in the U.S. in 2025 (+25% YoY; triple 2022 volume), with top-ten privacy settlements totaling ~$1.3 billion annually. This sits inside a broader class-action economy where settlements reached $79 billion in 2025.3 Because AI tools generate code carrying known security flaws in ~50% of builds4, unmonitored AI agents replicate these technical defects across 100% of user sessions continuously, scaling litigation exposure.5
AI Risk Selection and Evaluation: The Measurement Gap
Legal Alignment sits under Risk Selection and Evaluation: carriers taking AI risk must inquire into an AI system's safeguards and an organization's AI governance practices, and must conduct or require performance evaluations, red teaming, and chaos testing where applicable.6 Today, carriers can price a policyholder's disclosed controls, but they cannot yet price what an AI system does once no one is checking its output.
Other insurance lines have a comparable model. Property insurance uses catastrophe modeling. Liability lines track social inflation. AI liability has no actuarial history, and market agreement on what compliance looks like is still pending.
Standard ISO generative-AI exclusions, including CG 40 47 and CG 40 48, were published recently in response to this gap. Underwriters now choose between two weak options: write a blanket exclusion and give up a large market, or write specific policy covering an exposure nobody has priced.
Capabilities vs. Alignment: AI Is "Privacy-Blind" by Default
A measurement gap sits underneath this exposure. Traditional AI evaluations test legal capability: what a model knows on a quiz, a statutory research task, or a doctrinal question. These evaluations do not test legal alignment: whether an autonomous system's conduct follows statutory constraints while it runs.
Unprompted, an AI agent treats privacy as absent from the instruction "build me this site." A model can state a legal rule correctly on an exam, then ship code that violates that same rule, creating immediate class-action exposure.
Translating AI Conduct into Actuarial Evidence
To enable deliberate risk selection and accurate evaluation, insurers need real test data on how AI agents behave. DLAB is meant to bridge that gap, as the market’s first deterministic measurement layer for agentic compliance. The results yield five core takeaways for insurance professionals.
5 Key Takeaways for Insurance Professionals

#1 Some Models Deliver More Consistent Legal Alignment Than Others
Without instruction, all seven AI coding tools produce weak results. Mean protection of privacy is 27.8%. Privacy instructions change the mean significantly. General instructions give 63.5%. Expert specifications give 88.6%.
The best builds are close. The worst builds are far apart. Under the full specification, 20 of 30 Claude builds pass 95% protection, against 2 of 34 builds from the other tools. The best Claude build reaches 98.6%. The best build from another model family reaches 95.9%. The worst Claude build gives 83.6%, which equals the average of the other models. Their worst build gives 40.5%. The jailbreak instruction claims a lawyer approved the shortcut. In response, two Claude models add protection (Opus +16.7 points, Fable +8.7 points). GPT-5.6 removes protection (-11.4 points).7
The implications for insurance carriers underwriting companies deploying AI agents are clear. It's not enough to measure a policyholder's agents at their best build. Measuring the range between the worst results and the best is critical, because it will define severity of risk. On top of it, carriers should measure how a policyholder's agents respond when an employee says the legal review is complete.
#2 High-Level Governance Policies Fail; Operational Specifications Succeed
84% of developers use AI coding agents.8 To differentiate one policyholder from another, ask two questions:
- What specification do the agents build against?
- Does anyone test the running code for legal exposure?
A policyholder with a written implementation specification and an automated test of the running code sits in a different risk class from a policyholder with a governance policy alone.
DLAB measures the distance between such policyholders to narrow the field. Across all 37 parameters, that difference is worth 16.1 points of protection (between Self Research [72.5%] and Expert Guidance [88.6%]).9 On the parameters plaintiffs often sue over, it is worth 30.5 points.
Those parameters are:
- A working consent-withdrawl path that persists
- Pre-consent data leak
#3 Three Exposure Gaps, Ranked by Proximity to Litigated Precedent

DLAB reads the technical parameters against case law and statutory enforcement history. Three gaps carry real legal weight:
- Closest to litigated precedent — unconsented tracking (consent-withdrawal that stops nothing; pre-consent data leak): Agents built UI opt-out toggles that downstream trackers ignored (55 of 70 Bare Request builds failed this).10 This mirrors active California Invasion of Privacy Act (CIPA §631/§637.2) wiretap claims (Calhoun v. Google LLC, Javier v. Assurance IQ, LLC).11
- Regulatory enforcement risk — rights requests honored without checking who is asking: Agents created data deletion and access forms that executed without identity verification (52 of 70 Bare Request builds). This falls short of CCPA verification requirements and sits in the family of defects California has pursued through AG enforcement (Healthline Media, $1.55M12; Disney, $2.75M)."13
- Background risk multiplier — missing security headers: Absent CSP or HSTS headers do not create privacy exposures on their own but are the kind of gap a plaintiff's security expert may point to under CCPA §1798.150(a)(1)'s "reasonable security procedures" in standard breach litigation.14
#4 Pricing Exposure Against Real Class-Action Settlements
DLAB cross-referenced the first gap above — unconsented tracking–against Darrow's dataset of comparable common-fund wiretap settlements to establish expected-value loss curves based on class size.
These describe real operators with real class members; no build in this study has either. What they establish is what this class of defect has historically been worth when it appeared on a live site at scale.15

#5 Require Test Results, Not Just Specifications
Expert specification addresses all 37 parameters. But in some builds, the agent receives the instruction and does not follow it.
The shortfall concentrates in one place. A "Download My Data" request must return the user's actual records. 75% of builds return an acknowledgment instead or otherwise fail. This parameter is the worst for all seven configurations and the largest part of the residual.
Consent withdrawal shows the same pattern. A separate trial re-collected 50 specification builds with a working reject control.16 30% continued to load trackers after the user rejected.
These builds look correct from the outside. The rights page exists. The button renders. The specification behind them is complete. A reviewer who inspects the interface, the prompts, or the written policy finds no fault.
For underwriters, this holds a critical implication: a specification, a screenshot, and a policy document do not prove control. An automated test of the running code shows whether the system does what the legal specification says it should.

Conclusion
DLAB scores 404 builds from seven AI coding agents against 37 privacy parameters. Protection moves from 27.8% without instruction to 88.6% under a full specification. Exposure follows the technical guidance the agent receives. Guidance is a control a carrier can ask about, verify, and price. Darrow's wiretap settlement data attaches a dollar figure to the same defects by class size, from $0.5M at 35,000 class members to $115M at 220 million.
The specification does not execute itself. 75% of builds with the full specification fail to return a user's records from a working "Download My Data" button. 30% of builds with a working reject control continue to load trackers after the user rejects. These builds pass a document review and fail a test of the running code.
Three questions belong on the submission now:
- Which specification do your coding agents build against?
- When did you last test the running code, and what did the test return?
- What does your tooling do when an employee states that legal approved the work?
DLAB opens with a single domain and use case: privacy law for coding agents building websites. We will publish additional domains and use cases in the future.

.png)
.png)
.png)
.png)