DingDuff Benchmark

How a Frontier AI Model + DingDuff compares to the Legal-Specific Incumbent Models

We’ve given two frontier AI models, Claude and ChatGPT (with DingDuff access), the same hard veil-piercing assignment as Westlaw CoCounsel and Lexis Protégé, and then we’ve graded every memo in two ways: (1) substantive legal accuracy using an answer key we built by hand (applying TX and DE law), and (2) citation accuracy, checking all 382 citations against the actual source.

The upshot: Claude Fable with DingDuff scores a perfect 11/11 on legal accuracy and every citation is correct, and it outperforms the other models we’ve tested.

The full prompt we gave to every system

The Assignment

Please prepare a memo analyzing whether a trade creditor can pierce the veil of a Delaware LLC whose sole member is a Texas-resident individual. The LLC was formed in Delaware in 2019 to operate a single Houston-area restaurant. The sole member routinely paid personal expenses (his home mortgage, his wife’s vehicle lease, his children’s tuition) directly from the LLC operating account; the LLC never adopted anything beyond a one-page operating agreement, held no member meetings, and was initially capitalized with $5,000 against monthly operating expenses of roughly $80,000. My client, a produce wholesaler, is owed approximately $220,000 on open account. The LLC has ceased operations and is insolvent. Suit will be filed in Harris County. Please address: (1) whether Delaware or Texas law governs the veil-piercing analysis under Texas choice-of-law principles (internal affairs doctrine vs. substantive tort/contract characterization); (2) the substantive standards under each jurisdiction; (3) whether reverse veil-piercing is available; and (4) whether a companion Texas Uniform Fraudulent Transfer Act claim against the individual member is viable and how it interacts with the veil theory.

The Results

Frontier AI Models general-purpose

Fable 5

Anthropic · High Effort
11 / 11 correct
131 of 131 cites correct

ChatGPT 5.6

OpenAI · Extra High Effort
9 / 11 correct
50 of 51 cites correct
Legal-Research Tools commercial, purpose-built

Westlaw CoCounsel

Thomson Reuters · Extended
9 / 11 correct
87 of 91 cites correct

Lexis Protégé

LexisNexis
4 / 11 correct
68 of 89 cites correct
01

Substantive legal accuracy

11 sub-issues, graded right / wrong (or missed)

The prompt has four sub-questions, and we assessed accuracy based on the things that — in our own attorneys-who-have-practiced-in-this-area opinions — a correct answer would have to hit. We tried to focus on points that make a good binary (e.g. “did the AI find the controlling statute”), since the more intangible aspects of legal writing are hard to test for. Although if we were scoring on those softer factors, we’d also say Fable wrote the best memos, for what it’s worth.

The four main sub-issues are:

Q1  Which state’s law governs?
Q2  What’s the veil-piercing standard in each state?
Q3  Is reverse veil-piercing available?
Q4  Whether a companion TUFTA claim is also available.

Here’s our scorecard:

Legal Sub-issue
Frontier
Fable
Frontier
ChatGPT
Legal
Westlaw
Legal
Lexis
Q1 · Choice of law
1A
Delaware law governs
1B
§ 1.104 is the controlling statute
Q2 · Substantive veil-piercing standard
2A
Delaware: single economic entity + injustice
2B-i
Texas: § 21.223 is controlling
2B-ii
Texas: § 101.002 bridges it to LLCs
2B-iii
Texas: TUFTA actual fraud satisfies § 21.223(b)
Q3 · Reverse veil piercing
3A
Inapplicable here (forward, not reverse)
3B
Delaware recognizes it (Manichaean)
3C
§ 101.112(d) charging-order exclusiveonly Fable caught this →
Q4 · TUFTA companion claim
4A
Identifies controlling provisions
4B
Claim is viableLexis lucked in — never cited TUFTA →
11 possible points →
11
9
9
4
Legal Accuracy Notes

A few sub-issues deserve extra explanation for anyone reading the graded memos closely.

Q2B-ii
The corporation-vs-LLC trap. Westlaw AI missed this, but all other models got this. § 21.223 lives in the part of the Business Organizations Code that governs corporations, not LLCs. Not every corporate rule applies to LLCs, so an LLM that assumes it does isn’t analyzing deeply enough. A second statute, BOC § 101.002, applies § 21.223 to LLCs by reference — a clean test of whether the model is tracking the corporation/LLC distinction.
Q2B-iii
TUFTA / § 21.223(b) connection. Only Fable and Westlaw AI made this connection. Some cases recognize that actual-fraud asset transfers covered by TUFTA can satisfy the fraud requirement under § 21.223. We gave credit to models that spotted that case law, since this was a legal-research test.
Q3A
The biggest miss — BOC § 101.112(d). Texas statutorily foreclosed reverse veil-piercing (and similar remedies) for LLCs in 2023 with this emphatic bit of legislation. It arose after a man who owed his ex-wife $385k on a personal-injury judgment parked his assets in a wholly-owned LLC; the Fort Worth Court of Appeals allowed a cousin-remedy to reverse veil-piercing to stop the “I don’t own anything, but my LLC does” fiction. Our always-wise legislature responded, “never again.” The case law the models cited predates and is abrogated by this amendment. Only Fable High even flagged the statute — and even it framed the clash as an “unresolved collision” rather than controlling (which, in our opinion as Texas attorneys, it is).
Q4B
For the companion TUFTA claim, Lexis got to the right conclusion, but botched the reasoning. Because we only graded conclusions, Lexis gets credit — but it went way off the reservation. Its main source was a 1973 Delaware case cited to interpret a Texas statute passed in 1987, so the source has nothing to do with the statute. It wandered onto the right answer while addressing zero statutory provisions. We were genuinely surprised by how poorly it handled this.
02

Citation accuracy

every citation hand-checked
Correct cite source supports the point
Needs attention minor issue
Rejected wrong / misquoted / off topic
Not evaluated secondary sources excluded from %
Fable (High)
Frontier
134
131 evaluated, 131 (100%) correct
ChatGPT 5.6
Frontier
51
51 evaluated, 50 (98%) correct
Westlaw
Legal specific
108
91 evaluated, 87 (96%) correct; 17 not eval.
Lexis
Legal specific
89
76% correct · 10 need attention, 11 bad cites
How We Checked the Work

We read every source ourselves

We hand-reviewed every citation in all four memos (it took quite a while). We let Opus take the first pass at filling them in, which genuinely helped — it flagged errors we might have overlooked. Because we wanted anyone to be able to check our evaluations themselves rather than take our word for it, we excluded copyrighted secondary sources (treatises, Am. Jur., Restatements) that we couldn’t legally post online. Of the ones we did check, the models that cited secondary sources got them right.

The citation review panel explained

This is a tool we built to check work product before filing or use. It pairs the memo on the right with a text or PDF of the cited source (case, deposition transcript, statute) on the left. Click any citation and it pulls up that source. The highlights are an AI guess at the relevant passage — reliable for direct quotes, shakier on more complex points. We downloaded the case PDFs via Lexis (you can have Claude do it for you), but PDFs from Westlaw, Fastcase, or anywhere else work just as well.

If you want to run this on your own work product, it’s a free, open skill — point Claude at your memo and sources and it builds the panel for you.

Caveats

This is one prompt in one practice area. Substantive legal and citation accuracy reflect our own, necessarily subjective, opinion and judgment as the grading attorneys after reading the cases and statutes ourselves. “Not evaluated” citations were copyrighted secondary sources (e.g. treatises, Am. Jur., Restatements) that could not be posted online for public assessment of our scores — they are probably correct, but were excluded from the assessment rather than counted as errors. Pin cites and citation formatting were not evaluated, as not all systems produce pin cites or short cites.