How a Frontier AI Model + DingDuff compares to the Legal-Specific Incumbent Models
We’ve given two frontier AI models, Claude and ChatGPT (with DingDuff access), the same hard veil-piercing assignment as Westlaw CoCounsel and Lexis Protégé, and then we’ve graded every memo in two ways: (1) substantive legal accuracy using an answer key we built by hand (applying TX and DE law), and (2) citation accuracy, checking all 382 citations against the actual source.
The upshot: Claude Fable with DingDuff scores a perfect 11/11 on legal accuracy and every citation is correct, and it outperforms the other models we’ve tested.
The Assignment
Please prepare a memo analyzing whether a trade creditor can pierce the veil of a Delaware LLC whose sole member is a Texas-resident individual. The LLC was formed in Delaware in 2019 to operate a single Houston-area restaurant. The sole member routinely paid personal expenses (his home mortgage, his wife’s vehicle lease, his children’s tuition) directly from the LLC operating account; the LLC never adopted anything beyond a one-page operating agreement, held no member meetings, and was initially capitalized with $5,000 against monthly operating expenses of roughly $80,000. My client, a produce wholesaler, is owed approximately $220,000 on open account. The LLC has ceased operations and is insolvent. Suit will be filed in Harris County. Please address: (1) whether Delaware or Texas law governs the veil-piercing analysis under Texas choice-of-law principles (internal affairs doctrine vs. substantive tort/contract characterization); (2) the substantive standards under each jurisdiction; (3) whether reverse veil-piercing is available; and (4) whether a companion Texas Uniform Fraudulent Transfer Act claim against the individual member is viable and how it interacts with the veil theory.
The Results
ChatGPT 5.6
Westlaw CoCounsel
Substantive legal accuracy
11 sub-issues, graded right / wrong (or missed)
The prompt has four sub-questions, and we assessed accuracy based on the things that — in our own attorneys-who-have-practiced-in-this-area opinions — a correct answer would have to hit. We tried to focus on points that make a good binary (e.g. “did the AI find the controlling statute”), since the more intangible aspects of legal writing are hard to test for. Although if we were scoring on those softer factors, we’d also say Fable wrote the best memos, for what it’s worth.
The four main sub-issues are:
Here’s our scorecard:
A few sub-issues deserve extra explanation for anyone reading the graded memos closely.
Citation accuracy
every citation hand-checkedWe read every source ourselves
We hand-reviewed every citation in all four memos (it took quite a while). We let Opus take the first pass at filling them in, which genuinely helped — it flagged errors we might have overlooked. Because we wanted anyone to be able to check our evaluations themselves rather than take our word for it, we excluded copyrighted secondary sources (treatises, Am. Jur., Restatements) that we couldn’t legally post online. Of the ones we did check, the models that cited secondary sources got them right.
The citation review panel explained
This is a tool we built to check work product before filing or use. It pairs the memo on the right with a text or PDF of the cited source (case, deposition transcript, statute) on the left. Click any citation and it pulls up that source. The highlights are an AI guess at the relevant passage — reliable for direct quotes, shakier on more complex points. We downloaded the case PDFs via Lexis (you can have Claude do it for you), but PDFs from Westlaw, Fastcase, or anywhere else work just as well.
If you want to run this on your own work product, it’s a free, open skill — point Claude at your memo and sources and it builds the panel for you.
This is one prompt in one practice area. Substantive legal and citation accuracy reflect our own, necessarily subjective, opinion and judgment as the grading attorneys after reading the cases and statutes ourselves. “Not evaluated” citations were copyrighted secondary sources (e.g. treatises, Am. Jur., Restatements) that could not be posted online for public assessment of our scores — they are probably correct, but were excluded from the assessment rather than counted as errors. Pin cites and citation formatting were not evaluated, as not all systems produce pin cites or short cites.