AI for Lawyers, Companies, and Everyone Else

Draft

Version 0.14.1 · Effective 2026-08-31 · Last updated 2026-09-01

Last updated 2026-09-01 · v0.14.1 · draft · this document evolves, see Changelog

This guide is for people who use AI, work at companies that are figuring out AI, or practice law in a world where AI is already reshaping billable work. Skip to the part that fits you. Everyday users start at Part 2. Companies start at Part 3. Big Law starts at Part 4. Developers start at Part 5. Part 1 is background: read it if you want the landscape first, skip it if you want to get to work.


Part 1: The AI landscape in 2026

If you only read one part of this guide, read this one. It explains what these tools are, who makes them, what they cost, and how the ground shifted in the last six months. Everything after Part 1 assumes you have this picture in your head. I have kept the jargon out and the dates in.

1.1 What a language model is, plainly

Think of a language model as the most well-read intern you have ever met, one who has skimmed most of the public internet and recalls the shape of everything while recalling the details of almost nothing. You hand it the start of a thought, and it continues that thought the way a careful writer would. That continuation is the whole trick. Drafts, summaries, translations, contract markups, working code: all of it comes out of one move, repeated a few thousand times a second.

The model is good at a surprising range of work. It rewrites a clumsy paragraph into a clean one. It reads an 80-page lease and pulls out the ten clauses that could hurt you. It turns your half-formed idea into a first draft you can argue with. It explains a tax concept three different ways until one of them lands. It reasons step by step through a problem you describe, as long as the problem lives in words.

But the model cannot do several things people assume it can. It does not look anything up unless you connect it to a tool that searches. It does not remember your last conversation unless you turn on memory. It does not know today's news, your company's private files, or the contents of a document you have not shown it. And it does not actually know what it knows. Ask it a question it cannot answer, and it rarely says so. It writes a confident answer anyway.

That last trait is the one that catches people. The model picks the most plausible next words, not the true ones. When the facts are missing, it fills the gap with text that has the right shape: a citation that looks real, a statute number that parses, a statistic that sits in a believable range. The output reads as authoritative because reading authoritative is exactly what it was trained to do. The industry calls this hallucination, and it is not a bug the next release fixes. It is a side effect of how the thing works.

Even measured carefully, the productivity picture is unsettled. A 2025 randomized study of 16 experienced open-source developers completing 246 tasks in repositories they knew found that early-2025 AI tools made them 19% slower, although they expected to be faster.[98] METR's February 2026 follow-up then measured the same developers roughly 18% faster with late-2025 tools and newly recruited developers about 4% faster, with intervals wide enough to include zero.[99] METR calls that new data an unreliable signal, mainly because developers began declining to participate rather than work without AI, a bias it says pushes the estimate down, so the true speedup is probably larger.[99] Neither wave generalizes to every developer, model, or codebase. The durable finding is the gap between felt speed and measured speed. The lesson holds beyond code: the tool is genuinely useful and genuinely fallible, often in the same sentence.

So the practical question is never "is AI smart." The practical question is "is this specific task one where a fast, fluent, occasionally-wrong intern helps me or hurts me." The rest of this guide answers that question for three kinds of reader. Everyday users learn which tool to open and how to get more from it. Companies learn where the cheap, provable wins are. Lawyers learn where the malpractice traps hide. Builders learn how to wire these models into real software. Read on for the landscape, then jump to the part that fits you.

1.2 The five families that matter: OpenAI, Anthropic, Google, xAI, Meta

Five labs ship the models almost everyone uses. They trade the lead every few weeks, and no single model wins every task, so treat the rankings below as a snapshot from August 31, 2026, not a permanent order. Where benchmarks are contested, I say so rather than pick a winner.

The summer of 2026 changed the shape of this list more than the order of it. Every lab now ships a tiered family rather than a single flagship, and the tiers are priced far apart. You choose a capability level and a price at the same time, on the same model line. That is the single most useful thing to understand before reading the rest of this section.

OpenAI: the GPT-5.6 family, plus Codex as the work surface. OpenAI's current line is GPT-5.6, which went generally available on July 9, 2026 in three tiers: Sol at the top, Terra in the middle, and Luna at the bottom.[241] It replaced GPT-5.5 (April 23, 2026), which sat on a ladder running back through GPT-5.2 in December 2025 and GPT-5 in August 2025.[45][9] GPT-5.5 remains the reference point for OpenAI's strongest published job-task result: on the lab's own GDPval test it scored 84.9%, leading the field across the test's 44 real occupations.[8] The weakness has not changed, and it is the same weakness every lab has. The lead is narrow and the benchmarks come from the company that profits from them.

Pricing dropped sharply this summer, in two separate steps that are worth keeping apart. On July 30, 2026, OpenAI cut Terra by roughly a fifth and Luna by roughly 80%, to $2 and $12 and to $0.20 and $1.20. Sol was not part of that cut and held at $5 and $30. Sol came down separately on August 21, 2026, to $4 and $20. OpenAI says that promotional rate will remain available at least through November 21, 2026, but it has not announced the later price.[241] For requests above 272,000 input tokens, OpenAI charges twice the input rate and 1.5 times the output rate. Cache reads can cost one tenth of standard input, while cache writes can cost 1.25 times standard input. Batch pricing depends on model and endpoint, so check the current pricing page before treating any discount as universal.[241] In the chat app, Plus is about $20 a month, Go is $8, and Pro comes in two tiers: $100 a month for five times the limits and $200 for twenty times.[17][253]

The product shift is Codex: OpenAI now sells an agent workspace, not just a model, with desktop agents, ChatGPT workspace agents, scheduled work, approvals, admin controls, and analytics.[225][226][238] It also now sells people. On May 11, 2026, OpenAI launched the OpenAI Deployment Company, a professional services business backed by more than $4 billion from nineteen investment firms, consultancies, and systems integrators, and majority-owned by OpenAI. It embeds "forward deployed engineers" inside customer organizations, and it opened with roughly 150 of them through the acquisition of the consultancy Tomoro.[247] For a firm evaluating vendors, that matters: the model provider is now also a competitor to the systems integrator you might have hired to deploy it.

Anthropic: Claude Opus 5, with Fable 5 above it and Sonnet 5 below. Anthropic shipped four Claude 5 models in under two months, and the one most firms should default to is Claude Opus 5, released July 24, 2026.[239] Anthropic claims it lands within 0.5% of Fable 5's peak score on the CursorBench 3.2 coding benchmark at maximum effort, at half the cost per task, and that it more than doubles Opus 4.8 on Frontier-Bench v0.1 for the same price.[239] Those are the lab's own numbers on the lab's own chosen tests, so weigh them as claims rather than findings. Opus 5 runs thinking on by default and carries a 1M-token context window with 128K max output.[239]

Claude Fable 5, generally available June 9, 2026, remains the most capable model Anthropic sells widely, and it still leads Opus 5 on some narrow tasks.[223][239] Claude Sonnet 5 is the volume tier. Below both sits Claude Haiku 4.5 for cheap, fast, simple work.

The pricing tells you how to use the family. Fable 5 costs $10 per million input tokens and $50 per million output. Opus 5 and Opus 4.8 both cost $5 and $25. Sonnet 5 costs $2 and $10, and that price is now permanent: the increase to $3 and $15 that Anthropic had scheduled for September 1, 2026 was cancelled. Haiku 4.5 costs $1 and $5.[240] Opus 5 also exposes an effort setting (low through max) that trades tokens against thoroughness on the same model, which is usually a cheaper lever than dropping to a weaker model.[239]

The weakness across the family is still cost. Heavy agent use adds up fast, and quality can drift in very long sessions.

Google: Gemini 3.7 Flash, with Antigravity as the agent layer. Google's newest model is Gemini 3.7 Flash, which the company describes as its most intelligent workhorse model rather than its flagship.[242] The Pro line is still nominally the top tier, but 3.7 Flash outscores Gemini 3.1 Pro on several published indices while costing a fraction as much, so for most work it is the one to reach for.[242] Its genuine edge has not changed: one model reads text, images, video, audio, and PDFs together, which makes it the easiest choice when your material is not just words. It is also the cheapest frontier tier by a wide margin. Gemini 3.7 Flash and 3.6 Flash both run $0.75 per million input tokens and $3.75 per million output, but that is introductory pricing through December 31, 2026, and it rises to $1.50 and $7.50 on January 1, 2027.[242] Budget on the January number, not the current one. Gemini 3.1 Pro is still in preview at $2 and $12, and the older 3.5 Flash sits at $1.50 and $9.[242]

The caveat is unchanged: the much-cited reasoning scores are Google's own, measured Google's way, and the higher numbers you may see come from a separate, slower "Deep Think" mode.[25] Gemini 4 has been confirmed as in pre-training with no announced release date, specifications, or pricing, so treat any Gemini 4 claim you encounter as speculation.[242] Google's agent story runs through Antigravity and Managed Agents: one call can spin up an agent in an isolated Linux environment, with tools, code execution, and markdown-defined agent files.[236]

xAI: Grok 4.6. xAI's current flagship is Grok 4.6, released August 12, 2026 and aimed at long-running agents and interactive or visual work.[243] It costs $2 per million input tokens and $6 per million output, with a faster variant at double that.[243] It is best at three things: live data from X, fewer content guardrails than its rivals, and agentic work at a mid-range price. Its context window is reported at 500,000 tokens, the smallest of the current frontier models, which now sit at or above 1M.[243] For a lawyer feeding in a full production set, that ceiling matters more than the price does.

The predecessor Grok 4.3, which reached the public API on April 30, 2026 at $1.25 and $2.50, still holds the cheapest frontier output price and ranked first on Vals AI's Case Law v2 legal-reasoning benchmark at 79.3%.[202][203] That is a benchmark provider's own result rather than an independent courtroom test, and no published Case Law result yet covers Grok 4.6, so do not assume the legal- reasoning ranking carried forward. xAI's consumer pricing remains murky. Grok 4 launched in July 2025 alongside a $300-a-month SuperGrok Heavy tier.[47]

Meta: Muse Spark and Llama. Meta's consumer flagship is Meta AI powered by Muse Spark, announced April 8, 2026 as the first model from Meta Superintelligence Labs.[41] Meta also returned to genuine open source in August. Muse Glimmer, released August 10, 2026, is a 30-billion-parameter agent-optimized model under an unmodified Apache 2.0 license, which makes it meaningfully more permissive than anything in the Llama line.[244] A coding-focused Muse Spark 1.2 with a 1M-token context window is reported for August 5, 2026, and Meta has said its weights are "coming soon" without naming a date or a license, so treat Spark 1.2 as closed-weight until Meta says otherwise.[244] Llama 4 remains Meta's established open-weight family, the models you can download and run on your own hardware.[38]

Meta is best at two unglamorous things: a free assistant living inside WhatsApp, Instagram, and Facebook, and open weights for anyone who wants to self-host. The weaknesses are that public detail on the Muse line stays thin, and Llama is "open weight," not fully open source: a company with more than 700 million monthly active users needs a separate license.[219] For most people, both are free.

The pattern across all five is specialization, not dominance. Coding leads are often three or four points apart on the standard tests, and the same model can top one benchmark while trailing on the next.[26] Pick the model that fits the task in front of you, and re-check the ranking in a couple of months, because it will have moved.

1.3 Free vs paid: where the real lines are

Every major lab gives you a real free tier, and for a lot of everyday use the free tier is enough. You get a capable chat assistant, usually running a smaller or older model, with daily usage caps and limited or no access to web search, file uploads, memory, and connectors. If you ask a few questions a day and never hand the model your own documents, you may never need to pay.

You cross into paid territory when you need one of five things: the newest frontier model, higher usage limits, longer context for big documents, the ability to connect the model to your files and apps, or persistent memory across sessions. Most people who pay are buying limits and frontier quality, not a secret feature.

Two pricing models sit underneath everything, and confusing them wastes money. Per-seat pricing is a flat monthly fee per person, the way Netflix charges. It fits humans clicking buttons in a chat window. Per-use pricing, billed per million "tokens" (roughly, per few hundred thousand words in and out), fits software that calls the model automatically. A team of ten people drafting emails wants per-seat subscriptions. A nightly script that summarizes a thousand documents wants per-use API billing. Pick the wrong one and you either cap your people or torch your budget.

Consumer subscriptions, as of August 31, 2026:

Tool Free tier Paid
ChatGPT Yes, with limits Go $8/mo, Plus ~$20/mo, Pro $100/mo (5x limits) or $200/mo (20x)[17][253]
Claude Yes, with limits Pro $20/mo, Max $100/mo (5x) or $200/mo (20x)[184][253]
Gemini Yes, with limits US: Google AI Plus $9.99/mo, AI Pro $19.99/mo, AI Ultra $99.99 (5x) or $199.99 (20x)[27][29][253]
Grok Yes, inside X SuperGrok, up to ~$300/mo for the Heavy tier[47]
Meta AI Yes, inside Meta apps Free

Three of the four paid vendors have now converged on the same shape: a ~$20 everyday tier, then 5x and 20x usage multipliers at roughly $100 and $200.[253] That convergence is useful, because it means you can compare tiers across vendors on usage rather than on marketing names. These figures come from aggregator tracking rather than a single vendor page, so confirm the current number on the vendor's own pricing page before you commit a budget.

Business per-seat pricing, where the lines are sharper:

Product Per-seat price
Microsoft 365 Copilot Business $18/user/mo, paid yearly, first-year promotion through Dec. 31, 2026[164]
Microsoft 365 Copilot Enterprise $30/user/mo, annual[164]
Cursor (coding) Pro $20/mo[138]
Devin Desktop, formerly Windsurf (coding) Pro $20, Teams $80/mo base plus $40/full seat[171]
ChatGPT Business Standard $20/mo annually or $25/mo monthly[185]
ChatGPT Enterprise Quote-only, OpenAI does not publish it

For builders paying per use, the API token prices (per million tokens, input then output) tell the real cost story:

Model Input Output
Claude Fable 5 $10 $50[240]
Claude Opus 5 $5 $25[240]
GPT-5.6 Sol $4 $20 (promotional at least through Nov. 21, 2026, later rate unannounced)[241]
GPT-5.6 Terra $2 $12[241]
Claude Sonnet 5 $2 $10[240]
Grok 4.6 $2 $6[243]
Claude Haiku 4.5 $1 $5[240]
Gemini 3.7 Flash $0.75 $3.75 (introductory, rises Jan 1, 2027)[242]
GPT-5.6 Luna $0.20 $1.20[241]

Two numbers reframe the rest. The spread between the most and least expensive output in that table is more than forty to one, and the gap on the benchmarks between the top and the middle is a few points. Match the model to the job, and save the expensive model for the work that earns it.

Do not apply one provider's discount rules to another provider. OpenAI's GPT-5.6 Sol page, for example, distinguishes cache reads from more expensive cache writes and adds separate input and output surcharges above 272,000 input tokens. Google and xAI use different thresholds and storage rules. Anthropic offers cache and Batch discounts on eligible models, plus an effort setting on its newest models. Price the exact model, context length, cache operation, endpoint, and processing mode you will use.[239][240][241][242][243]

1.4 The three biggest shifts of the last six months

Three changes since late 2025 matter more than any single model release. The first changed how AI remembers. The second changed how it works. The third changed where AI work happens: from one chat box into operating surfaces that coordinate context, tools, agents, and model comparisons. All three started with practitioners, not press releases, and all three are simple enough to use this week.

Every major lab shipped a new model between this version and the last one, and most cut prices. None of that changed how anyone should work. The three shifts below did. When you are deciding what to pay attention to, that is the test worth applying: a new model changes what you can afford, but a new way of working changes what you do.

1.4.1 Knowledge bases (Karpathy's wiki): why it changed everything

The shift started with a post. On April 3, 2026, Andrej Karpathy described on X how he uses a language model to build and maintain a personal knowledge base, and the next day he published a short companion gist.[55][56] Developers seized on it within days. The idea is small and the payoff is large: instead of asking the model to fish facts out of a pile of documents on every question, you let the model build and keep a tidy wiki of its own.

The structure has three layers. A raw/ folder holds your original sources, untouched. A wiki/ folder holds clean pages the model writes and rewrites: one page per topic, per source, per open contradiction. A schema file (in Karpathy's setup, CLAUDE.md) tells the model what every page must contain and how the layers relate. I call this the functional triad: raw evidence, model-written knowledge, and a human-governed schema. Three operations keep it alive: ingest a new source, query the wiki to answer a question, and lint it to catch missing provenance, stale claims, and conflicts. Newer implementations make the provenance mechanical by requiring every wiki page to name its raw sources and by retaining disputed or outdated claims with status labels instead of silently replacing them.[286] Karpathy's own line on who does the work: "The LLM writes and maintains all of the data of the wiki. I rarely touch it directly."[128]

Here is what it fixes. The older approach, called RAG (retrieval-augmented generation), grabs raw chunks of your documents fresh on every question and forgets them the moment it answers. It is, in effect, amnesiac: it never learns, it just re-reads.[127] A maintained wiki does the opposite. Knowledge compounds. The model writes down what it figured out last week, reconciles a new source against the old ones, and starts the next question already smarter. For a personal or team-sized body of knowledge, plain markdown files plus a long-context model beat a complicated database.

This guide lives inside exactly such a setup. The repository you are reading from keeps source material under raw/, and the AI helps build structured notes alongside it. That is not a coincidence or a flex. It is the cheapest way I have found to keep a growing pile of research usable. You do not need to be a programmer to copy the pattern. A folder of plain text files, a chat model with file access, and three habits (add sources, ask questions, prune contradictions) get you most of the value. People report large wikis built this way on a single topic, though the exact sizes that circulate online are hard to verify, so treat those numbers with care.

1.4.2 Agentic loops (/goal, Ralph loops): why iteration beats one-shot

The second shift is about persistence. For most of the chatbot era, you asked one question and got one answer. An agentic loop changes the rhythm: the model takes a goal, does some work, checks its own result, and tries again, over and over, until the goal is met or it runs out of budget. Geoffrey Huntley named the bare-bones version of this pattern the "Ralph loop," a tight cycle where an agent re-reads its instructions and edits files using the filesystem as its memory.[62] The technique predates the products by more than a year.

In the last six months the labs baked it in. OpenAI shipped Codex CLI version 0.128.0 on April 30, 2026 with a built-in /goal command.[60] Under the hood it works through two template files, continuation.md and budget_limit.md: after each turn the agent is nudged to keep going toward the goal or to stop when it hits a token ceiling.[60] Simon Willison covered the release the day it landed.[58] Anthropic shipped its own official "ralph-wiggum" plugin for Claude Code in December 2025, folding Huntley's filesystem loop into the tool through its stop-hook system.[187]

What this makes possible is long-horizon work that no single prompt could finish: migrating a large codebase, repairing a broken test suite, or running an overnight build while you sleep. The model keeps its own progress notes in files and picks up where it left off.

But the failure modes are real and expensive. If the test the agent checks itself against is broken, /goal will happily burn 200,000 tokens trying to "fix the build" and get nowhere. Loops get stuck on a locked file, the wrong model, or context that has grown so long the model loses the thread. The better Ralph implementations now treat the loop as a small state machine: one backlog item, one fresh session, one objective check, one recorded state change, then repeat. They separate the task list, execution prompt, and learned guardrails, and they stop on repeated output or lack of progress.[284][285] That fresh context matters because the filesystem carries the durable state while each model run starts clean.

The rule that falls out of this: an agentic loop is only as good as its backpressure. Give it a reliable test, a bounded task, a stall detector, and a hard budget. A bad eval and a blank check let it spend money while going nowhere.

1.4.3 Agent operating surfaces: why one chat is no longer the unit

The third shift is that the unit of work is moving above the individual chat. A modern AI workflow may keep a project knowledge base, call tools, spawn scoped sub-agents, run browser actions, compare multiple models, and write a durable artifact at the end. The chat transcript is only one surface of the work.

The official product launches now make that visible. OpenAI's workspace agents are shared, Codex-powered agents that run in ChatGPT and Slack, can operate on schedules, ask for approvals, and expose analytics to admins.[226][238] Google's Managed Agents run the Antigravity harness in a cloud sandbox, with tools, code execution, and agent definitions stored in files.[236] Anthropic's Skills make the same point from the instruction side: reusable work belongs in a small, tested bundle with scripts and resources, not in a prompt you rewrite every time.[237]

Practitioner signals point in the same direction, but they should not be treated as product facts. OpenRouter's June 2026 description of Fusion shows one prompt fanning out to a panel of models while a judge model synthesizes consensus, contradictions, missing coverage, and unique insights.[206] Browser-skill products package repeatable web actions as reusable skills instead of bespoke prompts.[212] Neither product is the final form. Serious AI use is becoming a small operating system: memory, tools, model routing, review, and logs.

For buyers, the best answer may no longer come from the "best" single model. It may come from the workflow that gets three different models to disagree in useful ways, routes the mechanical work to cheap models, and reserves the expensive frontier model for judgment. It matters for lawyers because these surfaces create the audit trail the profession needs: what context was used, what tools ran, what models were consulted, and who approved the final output.

1.5 A 10-minute self-test: is this AI a fit for what I'm doing?

Before you reach for AI on a task, run it through this quick branch. Each branch ends with a verdict and points you to the part of the guide that goes deeper.

Is the task about writing, summarizing, explaining, or reshaping text I can give the model? If yes, the fit is strong. This is the core competence: drafting, editing, translating, turning long into short. Go to Part 2.

Does it depend on current facts, live prices, or recent news? If yes, the fit is conditional. Use a model with live web search and citations (Perplexity or Gemini with grounding), and verify anything that matters. A closed chat model will guess. Part 2.5 compares the research tools.

Does it need a specific fact stated exactly right, like a legal citation, a dosage, or a statistic I cannot check? If yes, the fit is poor without verification. This is the hallucination zone, and it is where the Mata v. Avianca sanctions came from. Use AI to draft, then confirm every hard fact against a primary source. Lawyers, see Part 4.7.

Is it a repeating business task across a team, like email, intake, or status updates? If yes, the fit is strong, but the win comes from setup, not magic. Build shared prompts and shared context once. Go to Part 3.

Does it involve my company's private documents or my clients' confidential data? If yes, the fit depends entirely on the tool's privacy terms. Use an enterprise tier with no training on your data, and read the contract. Part 2.7 covers the basics, and Part 4.6 covers the legal duties.

Am I trying to build software, automate a workflow, or wire a model into other systems? If yes, the fit is strong and the ceiling is high, but you are now a builder. If you are not sure you need code yet, read §3.6 first (it draws the line between staying in the chat UI and bringing in a developer). When you do need code, plan first, scaffold deliberately, and watch your token budget. Go to Part 5.

If none of these quite fit, the honest answer is to try it on a small, low-stakes version of the task and judge the result yourself. Ten minutes of testing beats ten minutes of theorizing.

And if you are a regular person who just wants to get more out of the chat window you already pay for, go to Part 2 next. Companies, Part 3. Lawyers, Part 4. Builders, Part 5. The landscape is behind you now. The rest is practice.


Part 2: The everyday user track

This track is a ladder, not a single step. You start by picking one chat app and using it daily (§2.1), add the workflows that earn back the subscription (§2.2–§2.5), then give it memory and your own files so it stops starting from scratch (§2.6). The last section (§2.12) shows how to graduate from plain chat to an agent that runs work for you, and points the way into the company and builder tracks when you outgrow chat entirely. Read it top to bottom the first time, then come back to the rung you need.

2.1 Pick your default chat

Pick one chat app, pay for it, and use it every day for a month. The habit matters more than the choice. All four major consumer chats are good enough now that your fluency with one beats the small quality gaps between them.

If you want a recommendation, here is the short version. Claude for writing and thinking through problems in prose. ChatGPT for general use and anything that touches your inbox. Gemini if you live in Gmail, Docs, and Android. Grok if you want live data from X and fewer guardrails. No single model wins every task, and that specialization is the defining feature of 2026.[133]

The prices cluster tightly at the entry level. ChatGPT Plus and Claude Pro both run about $20 a month, and Google's AI Pro is $19.99.[17][44][27] The free tiers are genuinely useful, so start there and upgrade only when you hit limits. The expensive tiers are for heavy users: ChatGPT Pro at $100 or $200 a month, Google AI Ultra at $99.99 or $199.99, and xAI's SuperGrok Heavy at $300, the priciest consumer tier among the majors.[17][29][47]

One recent change is worth knowing. OpenAI made GPT-5.5 Instant the default model in ChatGPT on May 5, 2026, and pitched it as less prone to making things up in areas like law and medicine.[46] If you tried ChatGPT a year ago and found it unreliable, the floor has moved.

The honest answer to "which one" is to try two for a week each and keep the one whose answers you trust. Fluency compounds. Someone who knows exactly how to ask their one tool will out-produce someone who keeps switching.

2.2 Five everyday workflows that pay for the subscription

Five workflows earn back a $20 subscription in the first week. None of them require any technical skill.

Drafting the email you keep putting off. The blank page is the expensive part. Paste the thread, say what outcome you want and what tone, and ask for a draft. You edit instead of compose, which is faster and less draining. ChatGPT and Gemini are convenient here because they sit close to your inbox.

Summarizing long documents you do not have time to read. Paste a 40-page lease, contract, or report and ask for the ten things that matter plus anything that looks like a risk. Claude suits this well because it handles very long inputs and produces long, structured outputs.[1] Always read the source for anything you will act on, but a good summary tells you where to look.

Turning a mess of notes into a clean draft. Meeting notes, voice memos, scattered bullet points. Hand the model the raw pile and ask for a structured version: agenda, decisions, action items. This is the workflow that quietly saves the most time per week.

Research you would otherwise farm out to a search engine and ten browser tabs. Ask a grounded research tool a real question and get a synthesized answer with sources. More on the specific tools in 2.5.

Explaining something until you actually understand it. A tax form, a medical result, a clause in a contract, a line of code. Ask for a plain explanation, then ask follow-ups until it clicks. The model never tires of the third "but why," and it does not judge you for asking.

The pattern across all five is the same. The AI removes the cost of starting. That is worth more than any single impressive output, because the tasks you avoid are usually the ones a subscription pays you back on.

2.3 Image generation: nano-banana, ChatGPT, when to use which

Two tools cover almost everything. Use Google's Nano Banana for a strong first image, and use ChatGPT when you need to change one thing at a time.

Google's image model, nicknamed Nano Banana (the current version is Nano Banana 2), became Gemini's default image generator on February 26, 2026.[26] It is the one to reach for when you want a polished result on the first try. It renders legible text inside images better than its rivals, handles photorealism well, and keeps a character looking consistent across a series of images.[26] Every image it makes carries an invisible SynthID watermark, so it can be identified as AI-generated later.[220]

It has a real limit, and Google says so. It can still misspell small text and get the numbers in an infographic wrong.[26] Check any image you will publish, especially anything with text or data in it.

ChatGPT's image tool is the other strong choice for iterative editing: keep an image and change exactly one thing ("same photo, make the jacket red") while the rest of the picture holds steady. Google built Nano Banana for the same kind of multi-turn editing, so as of mid-2026 both handle this job well. If your workflow is generate, react, adjust, adjust again, try each on your own images and keep whichever holds detail better for your case.

You do not pay separately for either. Image generation comes bundled in the subscriptions from 2.1: ChatGPT Plus at about $20 a month, and Google's AI tiers starting at $19.99.[17][27]

When to use which, in one line: Nano Banana for the first great image, ChatGPT to refine a specific one.

2.4 Voice modes and assistants

Voice is the mode most people try once and forget, which is a mistake. Talking to an AI beats typing for a whole class of tasks: thinking out loud, practicing a hard conversation, or getting an explanation while your hands are busy.

The two consumer voice modes worth knowing are ChatGPT's Advanced Voice and Google's Gemini Live. Both let you hold a natural back-and-forth, interrupt mid-sentence, and switch topics without starting over. They are included in the subscriptions covered above, so if you already pay for one, you already have a voice assistant.

The bigger shift is voice moving into the assistants you already use. On January 12, 2026, reporting confirmed an Apple deal for custom Gemini models to power a more personalized Siri later in the year, which would fold a frontier model into the assistant on hundreds of millions of phones.[218] Expect the line between "AI chat app" and "phone assistant" to keep blurring.

Two practical notes. Voice is best for input and thinking, less so for anything you need to read carefully afterward, so switch to text when you want to study the answer. And voice conversations are processed on the company's servers like any other input, so the privacy terms in 2.7 apply the same way.

Feature availability in this area changes fast, so confirm what your specific app and region support before you rely on a particular capability.

2.5 Deep research: Perplexity, Gemini, ChatGPT, Claude, Grok compared

Deep research is the mode where you ask one real question and the tool spends minutes, not seconds, reading many sources and writing a cited synthesis. It is the closest thing to handing a research assistant a brief. Five tools offer a version of it, and they differ in ways that matter.

Perplexity is the one built specifically for this. It leads with sources, attaches citations to almost every claim, and makes it easy to click through and verify. When your main worry is whether something is actually true, start here.

ChatGPT and Gemini both ship a deep-research mode inside their subscriptions. They tend to produce longer, more essay-like reports, and they are convenient because you are likely already paying for one. Gemini has an edge when your question touches Google's own data, like Maps or fresh web results.

Claude is strongest when the research is material you paste in yourself: a stack of documents, a long report, a set of transcripts. It reasons over what you provide better than it searches the live web.

Grok's distinguishing feature is real-time access to X, which makes it the choice for questions about breaking news or social reaction. It also posted the top score on one provider's legal-reasoning benchmark, 79.3% on Vals AI's Case Law v2, though that is the benchmark provider's own result, not an independent test.[203]

The honest caveat applies to all five. Every one can still cite a source that does not say what the tool claims it says. Deep research narrows the gap between AI confidence and reality, but it does not close it. Click the citations on anything that matters. The cross-check habit from 2.10 applies here too.

If you are building rather than chatting, the same research capability is now available as an API primitive: Gemini's URL Context tool and OpenAI's Responses API web_search tool let your own software ground answers on named pages or live search. The builder's view is in §5.9.[229]

2.6 Memory: what it is, why it matters, how to migrate it

Memory is the feature that lets an AI remember things about you between conversations, so you stop re-explaining yourself every time. It is genuinely useful and quietly risky, and most people never touch the settings.

ChatGPT introduced persistent memory on April 29, 2024.[11] It remembers facts you tell it and patterns it notices, then carries them into future chats. With the GPT-5.5 Instant update in May 2026, OpenAI added a way to see why the model remembered something, which makes the feature easier to trust.[46] Claude takes a different approach with Projects, where you load the relevant files and instructions into a workspace and everything in that project shares the context.[2] Gemini leans on memory sources and large file uploads, reading up to 1,500 pages of material you give it.[25]

Memory turns a generic tool into one that knows your writing style, your company's terminology, and your recurring tasks. The payoff compounds the more you use it.

The risk is the mirror image. Memory can be wrong, and it can be poisoned. A bad instruction saved to memory can quietly steer answers for weeks, and security researchers have shown that ChatGPT memory can be manipulated through ordinary-looking input.[52] Review your memory settings every so often and delete anything stale or odd.

On migrating memory between tools: no clean export-import button exists, and you should not wait for one. The practical method is to ask your current tool to summarize everything it knows about you and your preferences, then save that summary as a text file. Paste it into the new tool's memory or custom-instructions field. You own that file, it works across every model, and it doubles as a backup.

2.7 Privacy basics for non-paranoid people

The distinction that matters is account type: free and cheap consumer tiers may train on what you type, and paid business tiers generally do not. The settings differ by provider, so check yours. Anthropic, for example, updated its consumer terms on August 28, 2025: consumer Claude users now choose whether to allow their chats to be used for training, and the users who allow it have their data retained for up to five years, while those who decline keep the shorter 30-day window.[216]

Here is the rule that covers most people. Anything you would not want a stranger to read, do not paste into a free consumer chat. That includes client data, medical details, passwords, and unreleased work. For everyday questions, drafts, and learning, the consumer tiers are fine.

If you handle other people's confidential information, move up to a business tier with a no-training, low-retention agreement. OpenAI's ChatGPT Business and Enterprise plans commit to not training on your data by default. Zero data retention is a separate, request-only agreement rather than an automatic Enterprise feature, so ask for it explicitly if you need it.[21][20] Anthropic's Claude Team plan and Google's Gemini for Workspace offer comparable enterprise data controls.[167][168] The American Bar Association's Formal Opinion 512 tells lawyers to confirm exactly this before using any AI tool: check whether the provider trains on your inputs, and negotiate it away if needed.[83] That advice is sound for anyone with a duty of confidentiality, not just lawyers.

Two settings are worth finding on day one. First, the training opt-out toggle, usually under data controls, which on consumer tiers often lets you keep using the tool while declining to have your chats used for training. Second, memory, covered in 2.6, which decides what the tool retains about you across sessions.

If you are setting this up for a company rather than yourself, the same instinct scales into real controls: see §3.2 on connector permissions and §5.8 on the security surface of tools and MCPs.

The calm version of all this: match the tier to the sensitivity of the data. Casual questions on the cheap plan, confidential work on a business plan with a contract behind it. That single habit handles the large majority of real privacy risk.

One newer control is worth knowing about. OpenAI began rolling out a Lockdown Mode in ChatGPT on June 4, 2026, alongside Elevated Risk labels.[228] Lockdown Mode turns off the riskier surfaces, web browsing, image generation, Deep Research, Agent Mode, connectors, and downloads, to shrink the openings a prompt-injection attack can use. If you are handling sensitive material and do not need those features for a given session, it is the privacy-conscious default. The security mechanics behind it are in §5.8.

2.8 The meta-prompt trick (ask the AI to write your prompt)

Ask the AI to write the prompt, then run that prompt. This is one of the most underused moves in everyday AI use, though nobody has measured it against just asking. It works because the model has already read every prompt-engineering guide ever written. You haven't, and you shouldn't need to.

Here is the basic pattern: tell the AI what you are trying to accomplish and ask it to write a prompt you can give back to it, or to another AI, that will get you the best possible result. Then run that prompt. Practitioners sometimes call this "prompting the prompt."[287]

For consequential work, turn the trick into a prompt compiler. Ask the first model to interview you for missing facts, then write out one standalone prompt: what you want, what to work from, what is off the table, what the output should look like, how you will know it worked, and when to stop. Review that before you run it. Hand it to a fresh chat so it sees the finished instructions, not the back-and-forth that produced them. When a plausible mistake would be expensive, give the prompt and the finished work to a separate reviewer.

Nobody has published a study measuring whether the compiled prompt beats the direct request, and one early result cuts the other way for the newest models, which did better with a simple prompt than a heavily structured one.[310] I ran a small check of my own on a coding task: the compiler turned a vague request into instructions with five of the eight pieces I look for and surfaced seven questions I had not thought to answer, but it skipped the how-you-will-know-it-worked and when-to-stop pieces, and quietly added a step I did not want. The interview is where the value shows up. The review of what it writes is where you catch what it missed. Keep both.

A quick example. Say you want help with a lease dispute. You could type: "My landlord hasn't fixed my leaky faucet in three weeks. Help me write an email." That works fine. But try this first: "Write me a prompt I can give Claude to draft a polite but firm email to my landlord requesting urgent repair of a leaking faucet. The situation is: it's been three weeks, I notified him in writing on May 10th, and I want to preserve my options without burning the relationship." The AI will return a prompt that includes all the things a good prompt needs: tone guidance, relevant context slots, a request for a specific output length, and a framing that sets up the right register. Copy that prompt, fill in any remaining blanks, and run it.

A second example, for a company context: "Write me a prompt I can use to summarize a vendor contract and flag anything that could become a liability, written for a non-lawyer executive audience." The meta-prompt will prompt you to include the contract text, specify the audience, and ask the model to use plain language and flag concerns by risk level. That is a much better starting point than anything most people would draft from scratch.

Why does this work? Two reasons. First, the research on prompt structure is genuinely non-obvious: specific instruction placement, uncertainty protocols, output format specifications, and positive framing each shift quality measurably.[196] Most people writing prompts cold skip most of these. Second, the model can apply that knowledge to your specific situation in about two seconds.

When is this overkill? For simple, one-shot requests with no recurring use (ask it to explain a word, ask it to fix a typo), skip it. The meta-prompt pays off most when you will reuse a prompt, when the task has multiple interdependent requirements, or when you are building something that other people will use.

One practical note: the meta-prompt trick is especially good for building team prompt libraries. If your company is trying to standardize how it uses AI for a category of work (vendor summaries, client intake forms, status updates), running the meta-prompt and then iterating on the result is a fast way to get to a reusable template that actually works.

2.9 When the AI gets stuck: failure handling

Most people encounter AI failures and assume the tool just doesn't work for their use case. Usually, the problem is fixable in under a minute.

Refusal and safety blocks

When the model refuses a request, the first question to ask is whether the refusal is a genuine safety concern or a precision problem. Almost always, it is a precision problem. The model refused because your request was ambiguous, or it defaulted to a cautious interpretation.

The fix: add context and reframe positively. "Don't write anything that could be seen as threatening" becomes "Write this in a professional, measured tone that focuses on legal remedies." Specifying who you are and why you need something also helps: "I'm a nurse asking about medication dosages for patient education purposes" gives the model context to interpret the request correctly.

If you are genuinely hitting a policy block (adult content, certain weapons information, etc.), no amount of prompt rewriting will change that. Switch tools or rethink the task.

A related issue: Anthropic's prompt engineering guidance notes that Claude in particular can backfire on negative instructions.[197] Telling the model "don't make things up" may actually increase fabrication by making the unwanted behavior more salient. Say what you want instead: "Only state facts you are confident about. Use phrases like 'I'm not certain, but...' when you're unsure."

The hallucination spiral

Sometimes a conversation goes sideways and the model starts generating content that is increasingly wrong or increasingly confident about things it shouldn't be confident about. This is a context problem. The model has accumulated a conversation history that is steering it toward a bad pattern, and each new response builds on that accumulated wrong foundation.

The fix is to start a fresh conversation, not to argue with the current one. Paste in your original request with any clarifications you have developed from the failed conversation, and go again. Continuing to push in a broken conversation is the single most common way people waste AI time.

For longer tasks, this means building a habit of occasionally checking factual claims against fresh sources before the conversation goes too far in the wrong direction.

Tool and browsing failures

When a model with web browsing returns nothing useful, or returns a summary of a page that doesn't match what you know, the web tool probably fetched a paywalled version, a redirect, or a cached empty result. Options: paste the relevant text directly into your conversation, switch to a model with better web access (Perplexity is purpose-built for this, and is good at attribution), or just ask the model to tell you what it knows from training data and flag what needs verification.

Switching models

Different models have real, consistent differences for specific task types.

If you are doing research with real-time sources: Perplexity or Gemini 3.5 Flash with grounding are better than Claude or ChatGPT without web access. If you are doing multi-step reasoning over a document you have pasted in: the current frontier models (Claude Opus 4.8, GPT-5.5) are strong. If you are writing code: Claude and ChatGPT both handle this well, but they have different strengths across language and framework combinations. If the task is highly time-sensitive and needs current data: any model without live web access will struggle.

Switching is not a sign of failure. Running the same prompt across two models and comparing the results is a legitimate quality check.

2.10 Hallucination mitigation for everyday users

AI makes things up. This is not a bug that will be fixed in the next release. It is a structural feature of how these models work, and any practical guide to using AI has to spend time on it.

What hallucination is and why it happens

A language model generates text by predicting what comes next, token by token, based on patterns in training data. It is not retrieving facts from a database. It is not checking its claims against a source. When it generates a plausible- sounding citation, name, statistic, or URL, it is doing pattern-matching, not recall. The result looks authoritative and reads fluently, even when it is completely wrong.

Research shows that models have meaningful but imperfect self-knowledge: they can partly distinguish questions they can answer correctly from questions they cannot, though calibration is far from reliable.[188] The rest of the time they genuinely cannot tell, and they fill the gap with something plausible.

The high-risk task categories

Some categories of AI output carry much higher hallucination risk than others. Apply extra scrutiny to:

The cheap mitigations that actually work

You do not need to become a prompt engineer to reduce your hallucination exposure. These four habits cover most of the practical risk.

Ask for confidence levels. Adding "Tell me what you're confident about and what you're uncertain about" to any factual request shifts the model's behavior meaningfully. Research suggests the instruction works best with graduated language rather than a binary "say I don't know" command.[188] "I'm confident that... / I believe but am not certain that... / I don't have sufficient information" is more accurate than a flat flag.

Ask for sources, then verify them. "Cite the source for each claim" is worth adding to any prompt where you care about accuracy. Two caveats: if you paste in documents, asking for inline citations to those documents cuts hallucination substantially by grounding answers in text the model can point to.[200] If you don't provide documents and ask for citations anyway, the model may fabricate them. So the rule is: provide your sources, then ask the model to cite them. Don't ask the model to find sources it then has to invent.

Cross-check across two or more models. If you get a specific claim from one AI and it matters, run the same question through a second one. Agreement does not guarantee accuracy (both models may share training data or the same blind spot), but disagreement is a reliable flag that something needs human verification. For high-stakes work, do not merely ask "which answer is better?" Put the answers side by side and ask a third pass to extract consensus, contradictions, missing coverage, and unique insights. That is the practical version of the emerging multi-model fanout pattern: several models generate, then a judge or human reviewer compares.[206]

Use grounded research tools for factual claims. Perplexity and Gemini with web grounding are built to attach citations to live sources. For any claim that needs to be current and verifiable, these are better choices than a closed-context chat model. They hallucinate too, but less on current factual questions, and they show their sources.

What not to trust AI for without verification

The short list: specific legal citations, specific academic citations, specific statistics not drawn from a document you provided, any claim about recent events, and any domain where being wrong has real consequences for someone's health, legal situation, or money. AI is genuinely good at synthesizing, explaining, drafting, and reasoning. It is not a substitute for a verified primary source on specific facts.

The practical rule: if you would cite it to a client, a court, a doctor, or your boss, verify it.

2.11 Other high-value everyday techniques

These are techniques the research literature consistently supports but that most everyday AI users don't know about. Each one is low-effort and produces real improvements.

Use the universal prompt contract

For any task that matters, a good prompt is a small contract. It says what you want, what context matters, what good looks like, what evidence the model may use, what format you want back, and how the model should handle uncertainty. That is the durable pattern behind most prompt advice in this guide.

The everyday version is:

Help me [goal]. Context: [facts]. Audience: [who will read/use this]. Tone: [tone]. Format: [desired shape]. If anything is uncertain, say so instead of guessing.

If a prompt fails, do not immediately reach for a more elaborate trick. First ask which part of that contract is missing. Most bad outputs are missing audience, success criteria, evidence rules, or output shape.

Show examples of what you want (few-shot prompting)

The fastest way to get AI output in exactly the format and register you want is to show it one or two examples. This is called few-shot prompting in the research, and it is one of the best-supported techniques across all of the prompt engineering literature.[189]

Concretely: if you want a weekly status update in a specific format, paste in one or two previous updates you liked and say "write this week's update in the same format and tone." If you want a contract clause in a specific house style, paste in a clause from an existing contract as the example. The model uses the example primarily to lock in format and register, which is usually what you were actually struggling to specify in words.

Three good examples are better than ten mediocre ones. And the last example you provide has the most influence on the output, so put your best, most representative example at the end. One study found that example ordering alone could swing accuracy from near state-of-the-art to near random.[190]

Specify the output format

Tell the model exactly what you want back. "Give me this as a table with three columns: risk, likelihood, and mitigation." "Give me bullet points, no more than six, each under 15 words." "Return this as a JSON object with the fields name, date, and action_required."

Format specification is one of the core techniques the research identifies as explaining much of the variance in prompt quality.[196] Most people never use it and then wonder why the output is in the wrong shape. Adding one sentence of format instruction to your prompt costs nothing.

If you use Claude specifically: XML tags around sections of a complex prompt (like <context>, <examples>, <instructions>) help the model parse your input and produce structured output. Claude was trained to recognize this structure.[197] For ChatGPT, markdown headers work similarly.

Ask the model to grade its own output and rewrite it

After you get a first draft, add: "Now critique that response. What's weak, incomplete, or potentially wrong? Then rewrite it to fix those issues."

This two-step generate-then-refine loop produces measurably better outputs. One study found roughly a 20% average improvement across a range of tasks.[191] For code, the gains are especially large. Two caveats: limit this to one or two refinement passes (returns diminish sharply after that), and for pure factual recall, this technique can actually make things worse if the model "corrects" a right answer to a wrong one.[192] Use it for drafts, analysis, and code. Be more cautious with it for factual questions.

Put the important stuff at the start and end

Research from Stanford and UC Berkeley (Liu et al. 2023) found that AI models attend most to information at the beginning and end of a prompt, with roughly a 20 percentage-point accuracy drop for information buried in the middle.[102] This matters practically: if you have a long prompt with background context in the middle and your actual question at the end, the model may give your question less weight than the framing material.

The fix: state your main request at the top, put background context in the middle, and restate the core ask at the bottom. For long documents you are asking the model to analyze, the same logic applies: the most important passages are better placed near the start or end of your pasted text.

The recency-of-example effect

Related to the point above: when you give the model multiple examples, the last one carries disproportionate weight in shaping the output. This is the recency bias documented in Lu et al. 2022.[190] If you want the model to write in a particular style, or to follow a particular template, put your best and most representative example last.

On "you are an expert in X"

Role prompting is everywhere in AI advice. "You are a world-class copywriter." "You are an expert in tax law." The research consensus on frontier models is that generic expert labels have minimal effect on accuracy.[193] They reliably change tone, and they can introduce overconfidence without improving the underlying reasoning. The effect is domain-specific rather than uniformly absent: impersonating a relevant expert can shift results on narrow tasks.[194]

What works better: embed the methodology you actually want into the instruction. Instead of "You are an expert negotiator," try "Before suggesting a negotiating position, identify the other party's likely priorities and constraints, then suggest a position that acknowledges those while advancing my goals." You are describing a process, not a credential. When role prompting does help, the research finds the gain comes from activating a specific reasoning process, not from the credential itself.[195]

2.12 Graduating from chat to an agent

Most people stop at the chat box, and for a lot of work that is the right place to stop. But there is a ladder above it, and you do not have to be an engineer to climb the first few rungs. Each rung gives you something the one below it cannot, and each has a clear signal that you have outgrown it. Move up only when the signal shows up, not because the next rung sounds impressive.

Rung What it is What it gives you The signal you have outgrown it
1. Plain chat One chat app, used daily Drafting, summarizing, research, answers You paste the same context or run the same task every time
2. Chat with memory, projects, and connectors The same app, plus your files, memory, and tools (§2.6, §3.1–§3.2) Answers grounded in your documents and your accounts You want work to run without you watching every step
3. A consumer agent surface Claude Cowork or ChatGPT's workspace agents An agent that uses your apps, runs multi-step work, asks before acting It cannot reach a system you need, or you need real file and automation control
4. A coding agent Claude Code or Codex An agent that reads and writes files and automates multi-step work on your machine You are now building, see §3.6 and Part 5

Rung one is plain chat. It drafts, summarizes, and answers, and for occasional work that is enough. You have outgrown it the moment you notice you are pasting the same background or re-running the same task by hand every time.

Rung two is the same chat app with memory, projects, and connectors switched on, which §2.6 and §3.1 through §3.2 cover. Now the tool knows your writing, your terminology, and your files, and it can reach the accounts you connect. You have outgrown it when you want a job to run start to finish without you steering each step.

Rung three is a consumer agent surface. Claude Cowork is Anthropic's desktop agent that works across your files and documents, and ChatGPT's workspace agents play a similar role inside OpenAI's apps.[217] These run multi-step work, use the apps you connect, and pause to ask before they do anything consequential, and they do it without a line of code. You have outgrown this rung when the agent cannot reach a system you depend on, or when you need direct control over files and automation that a consumer surface will not give you.

Rung four is a coding agent. Claude Code and Codex run on your machine, read and write real files, and automate multi-step work, and the important thing for a non-technical reader is that you do not have to be an engineer to use them for file and automation tasks. You can point Claude Code at a folder of documents and ask it to rename, reorganize, or extract from them in plain English. The fastest way to learn the tool is to use it on itself: make a project called "Learning How to Use Claude Code," drop this guide into its project files, and ask the tool to build you a curriculum to become proficient.[235] §2.13 covers that move in full. The moment you are wiring a model into other systems or building something repeatable, you have crossed into building. Read §3.6 for the line between staying in the chat UI and bringing in a developer, then Part 5 for the builder track itself.

The ladder is not a race. Plenty of people get everything they need from rungs one and two, and that is a sound place to settle. Climb only when the work in front of you asks for it.

2.13 Learning the tool by using it

The fastest way to get good at a coding agent is to spend real usage learning it, not to read about it first. The move has four parts, and you can run all of them in an afternoon.

First, understand how your plan meters usage. Anthropic's paid plans measure use in a rolling five-hour session window, with separate weekly limits on top, and Claude Code draws from the same pool as the chat app.[234] The practical consequence is that you should treat a session as a budget to spend deliberately: run real tasks, watch where the usage goes, and you will quickly learn which kinds of work are cheap and which are expensive. Spending a window on purpose teaches you more about the tool's economics than any pricing page.

Second, make a dedicated learning project. Create a project called "Learning How to Use Claude Code," and put reference material into its project files so every session starts with context. Dropping this guide in is a reasonable starting point, because it gives the tool a shared vocabulary with you.

Third, meta-prompt your own curriculum. Ask the tool to build you a curriculum to become proficient, then work through it with the tool itself as the tutor. This is the meta-prompt trick from §2.8 pointed at your own learning: the model has read more documentation than you have time to, so let it sequence the lessons.

Fourth, read the official learning path. Anthropic maintains a learning hub at anthropic.com/learn that links its Academy courses for both users and builders.[235] Use it to fill the gaps your hands-on sessions expose.

The throughline: burn a window on real work, keep the context in a project, let the tool teach you, and check your understanding against the official path. That loop turns a week of poking into genuine fluency.

2.14 Health, money, and travel: the three places to slow down

The tools got dramatically better at your highest-stakes personal paperwork in 2026, and the evidence that they are safe to trust with it did not keep up. This section exists because the gap between those two sentences is where people get hurt.

Health. OpenAI launched ChatGPT Health in January 2026 and expanded it on July 23, 2026 to US users aged 18 or older on web and iOS, subject to rollout and product limits.[265] It can connect a hospital patient portal or Apple Health and answer plain-language questions about labs and visit notes. That is genuinely useful, and it is the single best use I know for these tools in a medical context: turning a discharge summary into English, building a question list before an appointment, and keeping a medication list straight.

Two cautions, both serious. In most consumer-app use, HIPAA protection does not follow records after they leave a covered provider. The result can differ when the app acts for a covered entity or business associate, so check the actual relationship and terms.[265] A coalition of state attorneys general also subpoenaed OpenAI in mid-2026 over its health-data handling.[265] And in July 2026 a Florida man sued OpenAI alleging that ChatGPT told him he had dysautonomia based on uploaded labs and imaging and advised him to stay recliner-bound, after which he suffered a pulmonary embolism that a physician attributed to the prolonged immobility.[266] The allegations are unproven, and I raise the case for the pattern rather than the verdict. The pattern is that a model will answer a diagnostic question confidently and will not tell you it should not have. Use these tools to prepare for a doctor and to understand what a doctor already told you. Do not use them to decide whether to see one.

Money. A 2026 test ran eight fictional tax scenarios through ChatGPT, Gemini, Claude, and Grok with the required forms supplied. Across the four chatbots and eight scenarios, the reported average absolute calculation error exceeded $2,000, and every model made at least one error.[267] Every needed input, still wrong. Tax preparation is exactly the kind of task these tools look competent at and are not: it rewards arithmetic precision and current statutory knowledge, and they are weak at both. Use AI to understand what a form is asking. Do not let it compute what you owe. And think twice before feeding a Social Security number and a full income picture into a chat window at all.

Travel and anything that spends money. Booking agents crossed from answering to acting in 2026. Google's Gemini Agent and Chrome's automatic browsing can compare and book across sites, and Delta's Concierge can cancel and rebook a disrupted itinerary on its own.[268] The failure mode changed with the capability: a wrong answer costs you nothing, and a wrong booking is a charge. One documented Delta case had the assistant tell a passenger her flight was safe, then cancel it and quote double to rebook.[268] Consumers have priced this in already. A June 2026 survey found 60% of UK consumers would abandon an AI shopping agent after a single mistake, and only 19% trust one to make everyday purchases correctly.[268] Let an agent assemble the options. Press the button yourself.

One thing to do this week, for your family. Voice cloning has made the grandparent scam cheap and convincing, and scammers now clone a relative's voice from a few seconds of social media audio to stage a fake emergency.[269] The countermeasure costs nothing: agree on a family safe word, and verify any urgent request for money by hanging up and calling back on a number you already have. Never the number that called you.


Part 3: The regular company track

This part is for the operator at a normal company. Not a tech firm, not a law firm, just a business with a CRM, an inbox, a calendar, and not enough hours. The goal here is return on a modest investment, proven fast, before you spend big.

3.1 The email, calendar, and CRM play

The highest-return move for most companies is connecting one AI assistant to the three systems where work actually happens: email, calendar, and CRM. Do that, and the assistant stops being a clever toy in a separate tab and starts handling the glue work that eats the week. This is a capability worth proposing to almost any team that sells, supports, or coordinates with people outside the building.

Pick the assistant that already lives where your data is. If the company runs on Google Workspace, Gemini connects natively to Gmail, Calendar, and Sheets, and Claude and ChatGPT both reach the same tools through their own Google connectors. If it runs on Microsoft 365, Copilot reaches Outlook, Teams, and Excel without extra wiring. The CRM connects on top: Salesforce and HubSpot both expose connectors the major assistants can read and write. The rule is to follow the stack you already have rather than adopt a new one for the assistant's sake.

The play, in order. First, connect the assistant to the inbox and calendar so it can draft replies, summarize long threads, and prep people before meetings. Second, connect it to the CRM so it can log activity, draft follow-ups tied to a specific deal, and surface which accounts have gone quiet. Third, set a standing instruction so the assistant runs a morning pass: what needs a reply today, which meetings need prep, which deals are stale.

Three workflows show the range.

The outreach tracker. Keep a single spreadsheet as the source of truth for every prospect: name, company, last touch, next step, status. Then let the assistant orchestrate it. New replies in the inbox get logged to the right row, the assistant drafts the next message in the sequence, and a daily pass flags who has gone cold and who is due for follow-up. One person can run outreach to dozens of clients this way without a heavyweight CRM, because the spreadsheet is the CRM and the agent is the operator.

The meeting-to-record handoff. A meeting ends, and the notes become a CRM update and a follow-up email without anyone retyping. The next time that account emails, the reply draft is waiting with the deal history already pulled in. This is the handoff that removes the most invisible work.

The morning triage. A standing instruction runs before the day starts: summarize overnight email, list what needs a reply, name the meetings that need prep, and flag stale deals. The person opens one digest instead of three apps.

Two stacks, gone deep. Google-native is Gemini for Workspace, or Claude or ChatGPT with their Google connectors, wired to Gmail, Google Calendar, and a Google Sheet as the tracker, with HubSpot added as the CRM once you outgrow the sheet. Glue the steps with the assistant's own scheduling, or add Zapier or Make for triggers the assistant cannot fire on its own. Microsoft-native is Microsoft 365 Copilot wired to Outlook, Teams, and Excel, with Dynamics or Salesforce as the CRM. Copilot's advantage is that it already sits inside the tools, so the connector work is minimal. The tradeoff is that it is strongest only when the whole company is on Microsoft.

Start with email and calendar alone for the first two weeks. They deliver value immediately and carry less risk than wiring an assistant into customer records. Add the CRM, or the spreadsheet tracker, once the inbox habit sticks.

3.2 Connectors, MCPs, plugins: explained without jargon

A connector is the bridge that lets your AI assistant read from and write to another app, like Slack, Notion, Salesforce, or Google Drive. That is the whole idea. The jargon around it (MCP, plugins, integrations) describes how the bridge is built, not what it does for you.

The term you will hear most is MCP, the Model Context Protocol. Anthropic published it as an open standard in late 2024, and it has become the common way for AI tools to talk to other software.[182] Think of it as a universal plug. Before MCP, every AI-to-app connection was custom. Now any compliant assistant can use any compliant connector, the way any USB-C cable fits any USB-C port.

The ecosystem exploded. Researchers cataloged 67,057 MCP server listings across six public registries using data collected in June and July 2025.[100] That is a listing count from a defined historical sample, not a count of live servers in 2026. The breadth is the appeal and the catch. You can connect almost anything, but not everything is safe to connect.

The security part is not optional, so here it is plainly. A connector is executable code that you are granting access to your data and, often, the ability to act on your behalf. The registry study found hundreds of vulnerable listings. Separate April 2026 research estimated that an affected MCP supply-chain pattern reached more than 7,000 public deployments, but code execution depended on downstream software accepting attacker-controlled STDIO configuration. The research did not establish that 7,000 hosts were successfully exploited.[100][137] The NSA published guidance on MCP security in May 2026, which tells you the stakes are no longer hypothetical.[111]

Four rules keep you safe without a security team. Install connectors only from sources you trust. Give each one the least access it needs to do its job. Require an approval step before any connector can create, change, delete, or pay for something. And pin to specific versions rather than auto-updating, so a compromised update cannot walk straight into your systems.[111]

3.3 Where to start: cheap pilots that prove ROI

Start narrow, on a high-volume task one team does every day, and measure before you scale. The data on AI pilots is sobering, and ignoring it is how companies waste money.

The sobering number first. A preliminary 2025 MIT NANDA report said 95% of the enterprise generative-AI initiatives it examined showed no rapid, measurable profit-and-loss impact.[106] The report drew on interviews, conference survey responses, and public initiatives rather than a representative audit of all enterprise pilots. Treat 95% as a directional finding from that sample, not a population failure rate. The initiatives that performed better targeted a specific, repetitive workflow rather than chasing a vague "adopt AI" mandate. The report also associated focused external tools with better outcomes than internal builds, but the evidence does not support a universal buy-over-build rule.[106]

So what does work? The evidence points to narrow, high-frequency tasks where time savings are easy to count. Anthropic's deployment with HUB International put Claude in front of 20,000 employees and, in vendor-published figures with no independent measurement, saved about 2.5 hours per employee per week on the targeted use cases.[160] A controlled study run with Boston Consulting Group consultants found that on tasks that suited the model, those using it completed 12.2% more tasks, did them 25.1% faster, and produced work rated over 40% higher in quality, while on tasks outside the model's strengths they did notably worse.[104] That last finding is the warning: match the pilot to what AI is actually good at.

Overreach is the other failure mode, and Klarna is the cautionary tale. The company replaced much of its customer support with AI, then publicly walked it back when the CEO conceded the work had come out at lower quality, and it began rehiring human agents.[54] The lesson is not that AI cannot do support. It is that scaling past what the model reliably handles, with no human backstop, costs more than the headcount it was meant to save.

Pick a pilot that meets four tests. The task is repetitive, it is done often, the output is easy to judge, and a human stays in the loop. Drafting first-pass email replies, summarizing long documents, turning meeting notes into structured records, answering routine internal questions. These are the proven starting points because they hit all four tests.

Two traps to avoid. The first is the productivity illusion: a METR study found experienced developers were actually 19% slower with early-2025 AI on mature codebases, yet believed they had been faster.[98] METR's 2026 follow-up then measured the same developers faster with newer tools, with intervals wide enough to include zero, and called its new data an unreliable signal of the current effect.[99] The trap survives either sign: measure real output, not the feeling of speed. The second is dead seats: a large share of paid Copilot licenses go unused within the first few months. Adoption below roughly a third of the team kills the return, so a small pilot with real usage beats a wide rollout nobody opens.

Once a pilot proves out, the next step is making the winning workflow run on a schedule instead of by hand. OpenAI's workspace agents, launched May 6, 2026 on Business, Enterprise, and Edu plans, are shared agents that run on a schedule, pause for an approval gate before acting, and report usage through an Agent Analytics view.[226] That combination, scheduled runs plus an approval gate plus analytics, is exactly the shape a repeatable workflow wants: deterministic, supervised, and measurable. Reach for it when a pilot has earned a permanent place, not before.

3.4 Building shared team context (project files, shared memory)

The difference between a team that gets real value from AI and one that does not is shared context. When every person re-explains the company from scratch in every chat, the tool stays generic. When the context lives in one place the whole team draws from, the answers get specific and they compound.

The cleanest pattern comes from Andrej Karpathy, who in early 2026 popularized treating an AI knowledge base like a wiki the model maintains.[55] You keep raw source material in one folder, the AI builds structured notes alongside it, and a plain instruction file tells the model how your organization works. Traditional document search forgets everything between sessions. A maintained knowledge base remembers, so each use adds to it instead of starting over.[127]

For a normal company, this looks like three things. A shared folder of source documents: policies, product specs, past proposals, the answers to questions people keep asking. A set of structured notes the AI helps keep current. And a short context file, the kind Anthropic calls a system prompt at the right altitude, that states who you are, how you talk to customers, and what good output looks like.[115]

3.4.1 The company context layer

For chat products, use Projects or their equivalent for recurring work. Put the stable files there once, then start fresh chats inside the project. Do not keep dropping the same files into unrelated chat windows. That creates version confusion and trains the team to use chat history as a filing cabinet. The better pattern is: stable context in project files, temporary reasoning in the chat, and a short handoff prompt whenever a new window takes over.

Assign owners to these files. Review them on a cadence. Archive stale drafts instead of leaving contradictions in place. The simplest rule is the best one: if three people have pasted the same explanation into AI, it belongs in shared context.

The team prompt library is the other half. When someone works out a prompt that reliably produces a good vendor summary or a clean status update, it goes in the shared library so nobody reinvents it. The meta-prompt trick from 2.8 is the fast way to build these: have the AI draft the reusable prompt, then refine it once with the team.

3.4.2 Prompt and skill libraries as operating infrastructure

Keep it boring and central. One folder, one context file, one prompt library, all in a place everyone can reach. Practitioner playbooks from the 2026 scrape converge on the same lesson: the instruction file should be short, explicit about conflicts, honest about token budgets, and paired with persistent notes rather than an ever-growing chat transcript.[208][209][213] The technology matters less than the habit of putting shared knowledge somewhere shared.

For teams, call this promptware rather than prompts. A reusable AI workflow should store the prompt, owner, model, required inputs, output format or schema, examples, known failure modes, review checklist, last-tested date, and version. That sounds heavier than "save this prompt in a doc," but it is the difference between a personal trick and operating infrastructure.

Skills are the next rung up from promptware. Anthropic's official guidance treats a skill as a discoverable bundle of instructions, scripts, and resources that Claude can load when the task calls for it.[237] That is the right mental model for a company: keep the active library small, name the owner, write a specific description, test it on real work, and retire stale skills. A 20-skill library that fires reliably beats a 200-skill drawer nobody trusts.

Do not dump the whole company brain into a skill. Use the skill for the procedure, and point it to the knowledge files it needs. A sales-call briefing skill might point to the CRM rules, customer-notes folder, approved tone guide, and output template. The skill says how to work. The knowledge files say what is true. One caveat from §5.5: the small slice of knowledge the assistant needs on every task belongs in the always-loaded project file itself, not behind a pointer it has to choose to follow, because measured evals show agents often never decide to look.

Give the library a promotion rhythm. If the team repeats the same workflow three times, write it down as a candidate skill. If the skill log has ten new patterns, review the cluster and decide what deserves to become shared infrastructure. Most skills should die as notes, not graduate. The point is not to collect skills. The point is to turn repeated work into a small set of reliable procedures.

For research-heavy work, separate raw source from synthesis. A useful structure is raw/ for untouched captures, notes/ for summaries, claims.md for important claims, contradictions.md for sources that disagree, open-questions.md for unresolved work, and synthesis.md for the current best answer. Never let the synthesis become the source.

3.5 Cost discipline: the per-seat math

Two pricing models exist, and confusing them is the most common budgeting mistake. Per-seat means a flat monthly fee per person. Per-use means you pay by the token for what you actually run through the API. Chat subscriptions are per-seat. Building on the API is per-use.

The per-seat field, as of August 31, 2026:

Tool Price per user / month
Microsoft 365 Copilot Business $18, paid yearly, first-year promotion through Dec. 31, 2026[164]
Microsoft 365 Copilot Enterprise $30 (annual, on E3/E5)[164]
ChatGPT Business (formerly Team) Standard $20 annually or $25 monthly, premium $100 annually or $125 monthly[185]
ChatGPT Enterprise Quote-only, OpenAI does not publish it
Claude Team $30, minimum 5 seats[167]
Google Gemini for Workspace starts ~$20, enterprise data controls[168]
GitHub Copilot Pro / Pro+ $10 / $39[166]

The per-seat math is simple. Multiply the seat price by the number of people who will genuinely use it, not the headcount you hope will. A large share of paid seats go unused within the first few months, so the realistic cost is the seat price divided by your true adoption rate. A $30 seat at 50% adoption costs $60 per actual user.

Per-use pricing matters once you build, and it fell hard over the summer. The frontier APIs as of August 31, 2026: Claude Opus 5 at $5 per million input tokens and $25 output, GPT-5.6 Sol at $4 and $20, GPT-5.6 Terra at $2 and $12, Claude Sonnet 5 at $2 and $10, and Gemini 3.7 Flash at $0.75 and $3.75 on introductory pricing that rises January 1, 2027.[240][241][242] Two of those are temporary. OpenAI cut Terra and Luna permanently on July 30, 2026, but Sol's $4 and $20 is promotional at least through November 21, 2026, and Gemini's rate doubles on January 1, 2027.[241][242] OpenAI has not announced Sol's later price. Test a conservative scenario in your budget, but label it as your assumption rather than vendor pricing. For a high-volume automated workflow, per-use can be far cheaper than buying everyone a seat. For interactive daily work, per-seat is simpler and usually wins.

One planning note that follows from the price cuts. If you built a cost model before July 2026 and rejected an automation as too expensive, re-run the arithmetic. A workflow that cost $30 per million output tokens on the old flagship now costs $12 on a mid-tier model that did not exist in June.

The discipline that saves the most money is not picking the cheapest tool. It is killing dead seats and matching the pricing model to the workload.

3.6 When to involve a developer vs stay in the chat UI

Stay in the chat app as long as a person is in the loop on every task. Bring in a developer the moment you want the work to run unattended, at volume, or wired into your own systems. That single line decides most cases.

The chat UI is the right tool for an enormous amount of work: drafting, summarizing, research, analysis, answering questions, one-off automations through built-in connectors. If a human is prompting and reviewing each result, you do not need to build anything. Most companies never need to cross this line, and crossing it early is a common way to overspend.

Four signals say it is time to build. You want a workflow to run on a schedule with nobody watching, like processing every overnight order. You need it to act across many records at once, not one at a time. You want it integrated with an internal database or a SaaS tool through a connector you control. Or the volume is high enough that per-use API pricing beats paying for seats.

There is now a middle tier between "stay in the chat" and "hire a developer," and it is worth knowing before you commission custom software. OpenAI's workspace agents let you stand up scheduled, approval-gated automations inside ChatGPT without writing application code,[226] and the May 2026 release notes add the operational details a company should care about: app-specific action safeguards, admin visibility, Codex goal mode, browser improvements, and usage analytics.[238] Its Apps SDK, previewed November 13, 2025, lets developers build interactive apps that run inside ChatGPT itself.[227] The Codex app pushes the same idea on the building side: a desktop surface for supervising several coding agents at once rather than one terminal session.[225] The practical read: try the no-code or low-code middle tier before you pay for a from-scratch build, and only cross into custom software when the middle tier genuinely cannot reach your systems.

The capability is real now, which is why the line matters. OpenAI's Codex added a persistent goal command in April 2026 that lets an agent keep working toward an objective across many steps, and Anthropic's Claude Opus 4.8 can plan a job and run hundreds of parallel sub-agents to do things like migrate a codebase of 100,000 lines.[60][1] This is genuinely useful and genuinely able to burn money. One team running parallel coding agents reported spending $7,000 in a single day.[138]

So the developer's first job is not writing code. It is putting guardrails on the autonomy: budget limits, approval gates, and a clear stop condition. Bring them in when you cross the line, and have them build the brakes before the engine.

3.7 Vendor selection (the short list per use case)

No single vendor wins everything, so pick by the job in front of you. The honest summary of 2026 is that the models specialize, and your existing software stack should break most ties.[133]

By use case:

By existing stack, which is the tiebreaker that matters most:

The rule that prevents most buyer's remorse: weight the vendor that integrates with the tools your team already opens every day. A slightly weaker model your people actually use beats a slightly stronger one they have to switch tabs to reach.

3.8 Evaluation harnesses and reproducible AI work

The next maturity step after "people are using AI" is "we can tell whether it is getting better." That requires a small evaluation harness: representative tasks, expected outputs or rubrics, model and prompt versions, the documents supplied, cost, latency, and reviewer notes. Without that, every vendor demo and every prompt tweak is judged by vibes.

Open-source diagnostics like iFixAi are useful less because any one score is definitive and more because they model the right shape of evidence: repeatable runs, scorecards, cross-provider judging, and content-addressed manifests.[210] For a company, start smaller. Save ten real examples, strip secrets, write down what good looks like, and rerun the same set whenever you change models, prompts, tools, or retrieval. If the new setup is faster but misses one of the ten things that matters, you learned something before it reached customers.

Rerun the eval set after five events: a model upgrade, a prompt edit, a tool change, a retrieval or context change, or a new high-stakes use case. Record the model, prompt version, tool permissions, context source, reviewer, cost, latency, and rollback rule. For legal or compliance workflows, add fixtures for citation accuracy, privilege/confidentiality handling, matter access, and approval logs.

3.9 From prompts to workflow design

Prompt engineering is no longer just clever wording. In 2026 the durable gains come from workflow design: what context gets loaded, what model handles which step, which tools are allowed, where the output lands, who reviews it, and how the result is stored for next time.

This is why prompt libraries should not only store text. They should store the whole pattern: source inputs, model choice, expected artifact, review checklist, and failure modes. A good "prompt" for vendor review or board prep is really a small operating procedure. The words matter, but the surrounding workflow matters more.

Here is what that looks like in practice. Take a recurring task most companies have: turning a sales call into a follow-up. The prompt-only version is "summarize this transcript and draft a follow-up email." The workflow version names every moving part. The input is the call transcript pulled from your meeting tool. A cheap model does the first pass: extract the attendees, the commitments made, and the open questions. A stronger model drafts the follow-up email in your house style, loaded from a project file. The output lands as a draft in your CRM, not in a chat window you will lose. A human approves it before it sends, and the approved version is saved back to the prompt library as the new template. The wording of the summarization step barely changes month to month. The decisions around it, which model, which context, where the result goes, and who signs off, are the part worth designing once and reusing. Section 3.8's eval harness is how you check that the workflow still produces the result you want after a model or prompt changes.

3.10 What actually goes wrong

Most of Part 3 tells you what to do. This section tells you what the evidence says happens instead, because a plan built only on vendor case studies is a plan built on a filtered sample.

Most projects do not deliver. Gartner surveyed 782 infrastructure and operations leaders in late 2025 about AI projects in their own infrastructure and operations function. In that sample, 28% delivered the promised return, 20% failed outright, and 57% of organizations reported at least one AI failure. The two leading causes were skill gaps and poor data quality, tied at 38%.[270] Those results do not describe every enterprise AI project. Read the two causes carefully, because they are worse at small-company scale, not better. In July 2026 BCG reported that nearly nine in ten CEOs saw some benefit in targeted areas while most were unable to scale it across the enterprise.[271] The pilot working is not the hard part. Getting past the pilot is.

Seats go unused, and the waste is larger than the sticker price. Industry tracking in 2026 puts weekly active use of paid Microsoft Copilot seats at roughly 20% to 30%.[272] Section 3.5's arithmetic applies directly: at 25% adoption, an $18 seat costs $72 per person actually using it. Buy in small batches, measure weekly active use, and reclaim what nobody touches.

Replacing people with AI has an unusually high regret rate. Reporting on Forrester's 2026 predictions said 55% of employers that cut headcount citing AI would later regret it.[273] Separate rehiring and restaffing-cost figures have been attributed to a different survey, but I could not obtain its primary report, so I have omitted them. The clearest named example is Commonwealth Bank of Australia, which cut 45 call-center roles on the claim that a voice bot had reduced call volume by 2,000 per week. Call volumes then rose, the bank apologized for the error, and it moved to rehire.[274] The lesson is not that automation never works. It is that the self-reported metric justifying the cut is often the thing that is wrong.

Agents connected to your systems are a security exposure, not just a productivity tool. A January 2026 online survey of 418 IT and security professionals reported that 65% had experienced at least one AI-agent-related incident during the prior year, with 61% reporting data exposure. Token Security financed the survey and co-developed its questionnaire with the Cloud Security Alliance, so treat it as vendor-commissioned evidence rather than a neutral incident census.[275]

The concrete case worth understanding is CVE-2025-32711, or EchoLeak. Security researchers demonstrated a zero-click attack path in Microsoft 365 Copilot: a later user query caused Copilot to retrieve a crafted email, follow its hidden instructions, and expose data available through connected Microsoft services. Microsoft patched the flaw before public disclosure and reported no known customer impact.[276] That is precisely the configuration Part 3 recommends most often, an assistant with access to mail and files, which is why §5.8's structural rule belongs in a non-technical company too. An assistant that reads untrusted input should not also hold broad, irreversible permissions.

None of this argues against the pilots in §3.3. It argues for the shape those pilots already have: small, measured, reversible, and judged on a metric you verified yourself rather than one the vendor supplied.


Part 4: The large law firm track

This part is for the large firm, written with the firm's risk posture in mind. Nothing here is legal advice, and none of it substitutes for your firm's own ethics counsel or your reading of the primary authorities. It is a practical map of what the market is doing and where the duties bite.

4.1 Why generic AI fails Big Law

A consumer chatbot is the wrong tool for a law firm, and the reason is the duty of confidentiality, not the quality of the model. The model is often excellent. The terms of service are the problem.

Free and consumer tiers frequently retain user inputs and may use them to train future models. Putting privileged client information into a tool with those terms risks the kind of disclosure that Model Rule 1.6 exists to prevent. The American Bar Association said as much in Formal Opinion 512 (July 29, 2024), which directs a lawyer to assess, before using any generative tool, whether the provider's terms allow it to retain or train on what the lawyer submits.[83] That single question rules out most consumer offerings on its own.

Confidentiality is the threshold issue, but it is not the only one. Generic tools lack the things a firm needs to use AI responsibly: an audit trail of what was asked and answered, citation validation against a real legal database, and a contractual guarantee that data is neither retained nor used for training. The enterprise legal products exist precisely to supply those controls.

Frame all of this as malpractice exposure, because that is what a risk committee will call it. A confidentiality breach through a tool's training terms, a fabricated citation that reaches a filing, or unverified AI output relied on as fact are each a path to a malpractice claim or a bar grievance, not merely a quality problem. The controls in this Part exist to keep that exposure inside the bounds the rules of professional conduct already require.

The hallucination stakes compound the problem. A model that invents a plausible case citation is a nuisance in most settings and a sanctionable event in a court filing. Damien Charlotin's public database of AI-fabricated-citation cases catalogued more than 1,200 globally by early 2026 and well over 1,600 by midsummer, growing by several a day. Section 4.7 has the current figures and the caveats on them.[89][251] Section 4.7 covers that risk in detail.

A chatbot every peer firm already has is not a differentiator. If every associate at every firm runs the same model, it raises the floor for everyone. That observation, which echoes through the market commentary, is what drives the largest firms to build rather than buy.[69]

4.2 The Harvey question

Harvey is the default answer to "what legal AI should we look at," and for many firms it is a reasonable one. It is also the company whose position best illustrates where the market is unsettled.

The facts first. Harvey raised $200 million in March 2026 at an $11 billion valuation, bringing total funding above $1 billion.[67][68] Its own account at that time described more than 100,000 lawyers across roughly 1,300 organizations in some 60 countries, and the company's current published figures are substantially higher, on the order of 200,000 lawyers and 2,400 organizations.[67][279] Press reports in August 2026 put Harvey in talks at a valuation around $15.5 billion.[279] Treat that as reported rather than closed, the same way you should treat the Legora figure below. It does contract analysis, due diligence, drafting, and research through chained agents, and in March 2026 it added a no-code Agent Builder that lets a firm assemble its own multi-step pipelines without engineers.[67]

Harvey does not publish per-seat pricing. Analyst estimates from Sacra put the base near $1,200 per lawyer per month for mid-market deployments with annual commitments, falling well below that for large Am Law volume deals.[69] These are secondary estimates, so treat any specific number as illustrative rather than quoted. Verified figures require direct negotiation.

The unsettled part is the business model. Harvey began with its own proprietary legal model, and according to the analyst firm Sacra it later moved to a model-agnostic approach after frontier models outperformed that proprietary system on Harvey's own internal benchmark.[69] That characterization is Sacra's, not Harvey's, so weigh it accordingly. The strategic question Sacra raises is whether an orchestration layer on top of frontier APIs faces commoditization pressure as OpenAI, Anthropic, and Google build legal-specific products directly.[69] Above the Law put the same worry more bluntly in March 2026, calling the current crop of legal tools "essentially AI wrappers" and predicting consolidation.[80]

Harvey is no longer the only well-capitalized answer. Legora, the Swedish competitor, raised $550 million at a $5.55 billion valuation in March 2026 in a round led by Accel, and its published customer list includes Cleary Gottlieb, White & Case, Linklaters, Goodwin, Bird & Bird, and Dentons.[249] Press reports in August 2026 put Legora in early talks at a valuation above $10 billion, which would nearly double its worth in four months.[249] Treat the $10 billion figure as reported rather than closed. The relevant fact for a buyer is not the valuation, it is that a second vendor now has enough capital and enough Am Law logos to be a real alternative in a competitive process. Run one.

The practical takeaway for a firm: Harvey is a credible buy, especially for getting capability into lawyers' hands quickly. But the "buy" decision should account for the possibility that today's vendor landscape looks different in two years, and it should be made against at least one serious competing bid.

4.3 Build, buy, or wrap

Most firms should buy, a few should wrap, and almost none should build from scratch. The right answer turns on scale, differentiation, and how much engineering the firm can actually sustain.

The buy case is the strongest for the large majority. A specialist vendor has already solved the confidentiality terms, the audit trail, and the legal-database integration. The much-cited MIT study of enterprise AI pilots found that buying from specialized vendors succeeded far more often than building in-house, though that study labeled its own methodology preliminary and has drawn criticism, so it is a directional signal rather than proof.[106]

The three subsections below walk the build-and-wrap options in order of increasing cost and risk.

4.3.1 Wrappers around Claude/OpenAI APIs

A wrapper is the fastest route to a custom tool: the firm builds its own interface and workflow logic on top of a frontier model's API, paying by the token rather than by the seat. Token pricing as of August 31, 2026 runs $5 per million input tokens and $25 output for Claude Opus 5, and $4 and $20 for GPT-5.6 Sol, with mid-tier options at roughly half those rates.[240][241] Harvey itself is, in effect, a sophisticated wrapper.

The economics of wrapping improved over the summer, which changes the build calculus. A firm that priced a wrapper against a $1,200-per-lawyer-per-month vendor seat in the spring should re-run that comparison, because the token side of it got cheaper while the seat side did not.

The appeal is control over the workflow without the cost of training a model. The risks are vendor lock-in, the engineering burden of maintaining the wrapper, data residency compliance, and the same commoditization pressure that the vendors face. A wrapper makes sense when a firm has a specific, repeatable workflow worth the engineering and an enterprise API agreement with confidentiality terms in place.

4.3.2 Fine-tuning on firm corpus

Fine-tuning, adjusting a model's weights on the firm's own documents, is the option most often proposed and least often worth it for knowledge injection. The evidence is unflattering. Stanford's FineTuneBench found that commercial fine-tuning APIs averaged only about 37% accuracy at absorbing new knowledge and far less at updating existing knowledge.[101] For getting firm-specific facts into a model, retrieval (pointing the model at your documents at query time) generally beats fine-tuning.

History offers a caution too. BloombergGPT, a 50-billion-parameter model trained on financial data, was overtaken by general frontier models not long after it shipped.[51] Fine-tuning still has a place for tone and format, and an open-weight model fine-tuned on-premises gives a firm physical control over its data. But it demands real machine-learning operations expertise and tends to trail the frontier on quality.

4.3.3 Full in-house model

Building a full proprietary platform is viable only at the very top of the market, where scale and differentiation justify the spend. The clearest example is Kirkland & Ellis, discussed next, whose reported commitment runs to $500 million.[63] For nearly every other firm, a from-scratch build is a way to spend a fortune reproducing what a vendor already sells. The MIT pilot data, contested as it is, points the same direction: most in-house enterprise AI builds do not pay off.[106]

4.4 What Kirkland and peers are reportedly building

The largest firms are building proprietary platforms, and the spending is substantial. Kirkland & Ellis is the headline. Bloomberg Law and Reuters reported in May 2026 that the firm committed roughly $500 million over several years to an in-house AI platform it does not intend to sell or license, with more than $100 million in the first year.[63][64] Coverage describes a team of around 180 technology professionals drawing on input from 250 lawyers, with roughly 50 dedicated AI engineers engaged so far.[63][65] Kirkland's chair is quoted, in secondary coverage that traces to a paywalled original, to the effect that off-the-shelf tools raise the floor for everyone but that clients do not hire the firm for the floor.[66] Treat the exact wording as reported rather than verified.

The build turned concrete in June. On June 4, 2026, Kirkland and Palantir announced a multiyear partnership to build an exclusive, proprietary fund enterprise platform for private capital fundraising, covering fund documentation, side letter drafting, obligation tracking, closing commitments, and ongoing compliance, with access for the firm's Investment Funds Group of more than 1,000 lawyers.[250] That is the most instructive detail in the whole Kirkland story, and it is easy to miss. The firm did not build a general-purpose legal assistant. It picked one high-margin, document-heavy practice where its own institutional knowledge is the moat, and it built there first. A firm with less than $500 million to spend can still copy the strategy: pick the single workflow where your accumulated know-how beats a general model, and build only that.

Other elite firms have taken different routes. Freshfields announced a multi-year partnership with Anthropic in April 2026, deploying Claude to 5,700 employees and co-building tools.[76] A&O Shearman was an early enterprise adopter of Harvey and co-built the ContractMatrix tool with Microsoft and Harvey, announced December 21, 2023.[221] Latham & Watkins rolled Harvey out firmwide on August 11, 2025, to more than 3,600 attorneys, though the deployment figures in wider circulation are breadth-level reporting rather than firm disclosures.[222] Several other firms are reported to be building internally. The specifics there are thinner and should be verified against primary sources before relying on them.

The pattern is a split. A handful of the largest, most profitable firms build to differentiate. Everyone else buys or wraps. The dividing line is whether a firm's scale makes a proprietary build cheaper than the strategic cost of using the same tools as its competitors.

4.5 Bridging chatbots and agentic coding

The frontier of legal AI is agentic: tools that plan a multi-step task and execute it across many documents, rather than answering one prompt at a time. This is the bridge between the chatbot a lawyer types into and the autonomous systems that engineers build.

Two developments mark the shift. Frontier models can now plan a job and run hundreds of parallel sub-agents in a single session, which makes document-scale or matter-scale work tractable. Claude Opus 5, released July 24, 2026, supersedes the Opus 4.8 model this section previously described, and dynamic workflows made the parallel-agent pattern a standard feature rather than something you assemble yourself.[239][245] And Thomson Reuters rebuilt CoCounsel on Anthropic's Claude software development kit as an agentic product in April 2026, serving a reported one million users across 107 countries, then shipped a next-generation CoCounsel Legal in August 2026.[73][74][255]

The integration layer is maturing alongside the models. In May 2026 Thomson Reuters launched a CoCounsel Legal MCP server that connects Claude and other MCP clients to CoCounsel Legal and the roughly 1.9 billion documents in Westlaw, which means the agent reasons over authoritative legal sources rather than its training data.[231] That is the difference between a model that sounds like a lawyer and one wired into the materials a lawyer actually cites.

The open-source side stopped being a signal and became a product. When earlier versions of this guide described Anthropic's claude-for-legal repository, it was an early indicator that legal AI was being packaged as workflow artifacts rather than generic chat advice.[215] On May 12, 2026, Anthropic shipped it as Claude for Legal, with twelve practice-area plugins, named workflow agents, and more than twenty connectors into legal systems of record.[248] Section 4.10.1 covers what that changes. The short version for a buyer: ask vendors what repeatable workflow assets they provide, and ask what those assets do that an open-source plugin from the model provider does not.

The distinction that matters for buyers is between genuine agents and marketing. A real agent plans, acts, runs its own tools or checks, and iterates without a human at each step. Much of what is sold as "agentic" is closer to chat with autocomplete.[73] Ask a vendor to demonstrate unattended multi-step execution, not a single clever completion.

Where these loops fail is predictable and worth designing against: ambiguous goals, missing stop conditions, and the absence of a human approval gate before any irreversible action.[89] For a firm, the approval gate is not optional. No agent should file, send, or commit anything without a lawyer signing off.

4.6 Compliance, audit trails, IP, work-product doctrine

The governing framework for a US firm starts with ABA Formal Opinion 512 (July 29, 2024), which maps generative AI onto existing duties rather than inventing new ones.[83] It implicates competence (Rule 1.1), confidentiality (Rule 1.6), communication with the client (Rule 1.4), candor to the tribunal (Rule 3.3), supervision of lawyers and nonlawyers (Rules 5.1 and 5.3), and fees (Rule 1.5).[83][84] The opinion's confidentiality guidance is the most operationally demanding: it contemplates informed client consent, beyond boilerplate engagement language, before client information goes into a self-learning tool.[83]

State authority has proliferated alongside the ABA opinion. California issued practical guidance in November 2023, and the New York City Bar published an opinion.[85][86] Avoid unsourced counts of state opinions because trackers group formal opinions, guidance, and proposed rules differently.

California also moved toward binding law. As of August 31, 2026, Senate Bill 574 remained an active bill in the Assembly floor process and was not law. The official text had last been amended on August 21.[252] That text would prohibit an attorney from delegating the practice of law to generative AI. It would also require reasonable steps to verify AI output, including case and statutory citations, and to correct erroneous or hallucinated material used by the attorney. For documents submitted to a court, it would require disclosure of AI use and personal verification of every filed citation. Separate provisions would restrict entry of nonpublic information and an arbitrator's delegation of decision-making.[252]

The status label matters as much as the substance: pending, not law. Check the current status and enrolled text on the California Legislative Information site before relying on the bill, because a floor vote, amendment, or veto can change both the obligation and its effective date.[252]

Three operational requirements follow. First, an audit trail: the firm should be able to reconstruct what was asked, what was returned, and who reviewed it. Second, a documented reasonableness analysis of the vendor's terms under Rule 1.6(c). Zero retention and no training on firm data are strong controls and worth negotiating for, but Opinion 512 requires a reasonable-efforts analysis rather than any specific contract term, and many compliant enterprise agreements retain data briefly for abuse monitoring.[83] The dangerous inference runs the other way: zero-retention terms do not discharge the duty, because Opinion 512 separately contemplates informed client consent for self-learning tools, and no contract term substitutes for that. Third, a clear human review step before any AI output reaches a client or a court.

Two things that Part 4 previously named and did not explain deserve their own treatment, because both are live exposure rather than background doctrine.

Billing is the exposure nobody budgets for. Opinion 512 addresses fees under Rule 1.5, and the guidance is more restrictive than most firms assume. A lawyer billing hourly may bill only time actually spent, may not bill for the time AI saved, and generally may not bill a client for time spent learning the tool.[83][280] Whether AI costs pass through to the client has to be disclosed in advance. And a flat fee left unchanged after AI compresses the underlying work can itself become unreasonable.[280] Read that against §4.8's instruction to measure real time saved: the saved time is a margin question and a client communication, not billable inventory. Before a practice group adopts a tool, revise the fee agreements, decide the cost pass-through, and tell the billers. Do it in that order, and do it before the first invoice rather than after the first fee dispute. I would treat these as reported summaries of the opinion and read the fee section of Opinion 512 directly before setting firm policy.

Confidentiality and privilege are different problems. Rule 1.6 is an ethical duty. Attorney-client privilege and work product are evidentiary protections with separate elements and waiver rules. A disclosure can satisfy an ethics policy and still lose a protection.

United States v. Heppner, No. 1:25-cr-00503-JSR, ECF No. 27 (S.D.N.Y. Feb. 17, 2026), supplies the warning that earlier versions of this guide missed. The court held on its facts that a represented defendant's consumer-Claude material was neither privileged nor work product. The defendant had used Claude on his own initiative rather than at counsel's direction, and disclosure under the consumer service's terms defeated confidentiality and waived any otherwise applicable privilege.[281]

That holding does not settle counsel-directed use under a negotiated enterprise agreement. A firm might argue by analogy to United States v. Kovel, 296 F.2d 918 (2d Cir. 1961), but Kovel is Second Circuit authority, not a universal third-party exception. It protects confidential communications made to obtain legal advice when the nonlawyer is necessary, or at least highly useful, to help the lawyer render that advice. It does not protect a client's use of a third party's own service merely because a lawyer is involved.[281]

Enterprise confidentiality terms can help prove those elements, but they do not create privilege. Contract for no training, tightly limited human and subprocessor access, deletion rights, and a defined legal-services purpose. Document counsel's direction and why the tool was needed. Then assume the enterprise-AI application of Kovel remains open until controlling authority says otherwise.[281]

4.6.1 Audit trails should include the workflow, not just the final text

A useful audit trail records more than the answer. It should identify the model, prompt, documents supplied, tools called, intermediate outputs that mattered, human reviewer, and final approval. In an agentic workflow, the path is part of the work product. If the firm cannot reconstruct that path, it cannot supervise the system in the way Opinion 512 expects.

One warning belongs next to the requirement because the requirement creates it. An audit trail is electronically stored information, but that label does not make every log discoverable or subject to a hold. Preservation begins when litigation is reasonably anticipated and extends to relevant material. Discovery then depends on relevance, proportionality, possession, custody or control, and privilege or work-product protection under Rules 26, 34, and 37. A privilege log is used for responsive material withheld on a protection claim, not for every record that reflects judgment.[281]

Build the log because supervision may require it, and build governance around it at the same time: a retention schedule, defensible suspension of destruction when a hold applies, matter-scoped access controls, and a process for reviewing responsive material. Treat those as risk controls, not categorical discovery rules.

This is now a procurement criterion you can test, not just an aspiration. Thomson Reuters' HighQ MCP server, released May 19, 2026, is built read-only, permission-controlled, and fully audit-logged, so every agent access to matter data is scoped and recorded.[232] When you evaluate a legal-AI connector, ask whether it logs every access, enforces existing matter permissions, and defaults to read-only. Those three properties are what turn an audit trail from a promise into evidence.

4.6.2 Redaction and anonymization are workflow controls

For sensitive matters, confidentiality may require changing the workflow before the model ever sees the data. Vendor patterns such as local anonymization, metadata stripping, and file-level filtering before an agent receives access are useful examples of the control layer.[211] Treat them as controls to evaluate, not magic shields: test whether the redaction survives realistic documents, logs, attachments, filenames, and tool outputs.

Two harder questions remain genuinely open. Who owns the output of a generative tool, and whether a lawyer's proprietary prompt sequences submitted to an external model enjoy work-product protection or are discoverable, are both being worked out in real time. There is no settled answer yet, and a firm should treat both as live risks rather than resolved doctrine.

4.6.3 Prompt discovery depends on who used the model and why

The 2026 cases do not point in one direction. They split by user, purpose, confidentiality, and the protection asserted.[254][281]

In Conservation Law Foundation, Inc. v. Shell Oil Co., a magistrate judge in the District of Connecticut ordered a testifying expert to produce generative-AI prompts used to filter document productions. The order treated the filtering as part of the expert's methodology, not as attorney prompts generally. The court stayed the production order while a Rule 72(a) objection was reviewed, so the order should not be described as an operative final answer.[254]

Other courts protected different material. In Morgan v. V2X, Inc., the District of Colorado held that Rule 26(b)(3) could protect a self-represented litigant's AI-assisted litigation preparation and that use of a third-party AI service did not automatically waive work product. The court still required the litigant to identify the platform and amended the protective order to restrict confidential information to services with specified contractual safeguards. In Assini v. Hayward, a New York court adopted similar reasoning and quashed a subpoena seeking a self-represented litigant's OpenAI account materials.[254] Heppner, discussed above, reached the opposite result for a represented criminal defendant who used consumer Claude without counsel's direction.[281]

The operational rule is to classify the use before predicting protection: attorney or client, represented or self-represented, litigation preparation or ordinary research, expert methodology or draft analysis, consumer terms or a negotiated enterprise agreement. Then apply the governing jurisdiction's actual privilege, work-product, expert-discovery, and protective-order rules.

Revise expert engagement letters now. State whether the expert may use AI, for what purpose, what must be retained, and what may be discoverable as methodology. For counsel and clients, record the legal purpose, terms, access controls, and direction of use. Do not delete a prompt history because it became awkward after a dispute. Apply the ordinary preservation analysis first.

The canonical cautionary tale is Mata v. Avianca, No. 22-cv-1461 (S.D.N.Y. June 22, 2023), in which counsel submitted a brief citing cases that did not exist, generated by ChatGPT. The court imposed a $5,000 monetary sanction jointly and severally against the two attorneys and their firm, payable to the court registry, and courts in the aftermath adopted standing orders requiring attorneys to certify they have verified AI-assisted content.[88][92] Describe Mata as a sanctions matter, because that is what the public record supports. Do not overstate it as a doctrinal holding on AI.

The problem did not stay contained to one case, and it is accelerating rather than fading. The Charlotin database tracked more than 1,200 fabricated-citation incidents worldwide by early 2026, and secondary trackers citing it reported figures in the 1,500 to 1,700 range by midsummer 2026, growing by several entries a day.[89][251] Norton Rose Fulbright, reviewing the 2026 record, counted more than 1,148 documented US attorney-hallucination cases.[251] I have not been able to pull a current total directly from the database, so treat the higher numbers as reported rather than verified, and check the database itself for today's count.

The size of the penalties moved too, but aggregate headlines hide important differences. In Couvrette v. Wisnovsky, the District of Oregon entered a $15,500 sanction on December 12, 2025 for fifteen fabricated citations and eight false quotations. A separate March 23, 2026 order awarded $94,704.38 in fees and costs. The combined $110,204.38 spans two orders, two remedies, and two calendar periods.[258] I could not reproduce a reported $145,000 first-quarter aggregate from primary orders, so this version does not repeat it.

The other Oregon authority is Doiban v. Oregon Liquor & Cannabis Commission, 347 Or. App. 742, 747–48 (2026), which imposed $10,000 for fifteen false citations and nine contrived quotations.[258] Earlier versions of this guide wrongly called that matter In re Ghiorso. William L. Ghiorso was counsel in Doiban. In re Ghiorso is a separate 2016 disciplinary matter. The correction belongs in the text because mismatching a real case name with a real sanction is exactly the failure this section warns about.

One finding from that review should reassure anyone waiting for a rule change to tell them what to do. Courts have not needed a generative-AI-specific rule. Existing Federal Rules of Appellate Procedure, Federal Rules of Civil Procedure, and state professional conduct rules have proved sufficient to sanction this conduct.[251] The duty you are already under is the duty that applies. Reported examples since Mata include a Sixth Circuit order in Whiting v. City of Athens (March 2026) imposing $15,000 per attorney plus opposing counsel's appellate fees for more than two dozen fake citations across three briefs, plus double costs and a disciplinary referral.[91][201][251] Whiting carries the same caveat this guide applies to Mata: reporting indicates the Sixth Circuit did not expressly attribute the fabrications to generative AI. It sanctioned the filing of unverified citations however they were produced. Calling it an AI case overstates the holding.[251] The 2026 record also includes a $2,500 fine in Fletcher v. Experian (Fifth Circuit, February 18, 2026), public admonishment in In re: Nwaubani (Fourth Circuit, March 11, 2026), and a public reprimand in Fivehouse v. DOD (E.D.N.C., April 27, 2026).[251] Not every case draws a sanction: in Gamez v. County of Fresno (E.D. Cal., April 9, 2026) the court imposed none.[251] The range runs from nothing to career damage, and which end you land on depends largely on whether you caught it first and told the court. One federal court applied Rule 11 sanctions for AI citation failures in Bunce v. Visual Technology Innovations, No. 2:23-cv-01740 (E.D. Pa.), where the court entered its sanctions memorandum and order on April 20, 2026, imposing $5,000 and mandatory AI-ethics continuing legal education on the sanctioned attorney.[278]

A word about that last citation, because it is the point of this section. An earlier version of this guide cited Bunce as "2025 WL 4231632 (E.D. Pa. Jan. 21, 2025)." That citation was wrong. The Westlaw number does not correspond to the opinion, and the date preceded the actual ruling by fifteen months. It came into the draft from a secondary source and survived several review passes, inside the section warning readers about exactly this failure. I caught it only when an adversarial review pulled the docket. Cases reported only through secondary sources must be verified against the docket before a firm cites them, and I am the proof rather than the exception.

Two lessons carry across all of these. Rule 11 responsibility attaches to the lawyer who signs, regardless of which tool drafted the text, so the duty to verify cannot be delegated to software. And using one AI to check another is not independent verification: in Mata itself, counsel asked ChatGPT to confirm the cases were real and it falsely affirmed them, which is exactly how the fabricated citations reached the brief.[88]

Checking that the case exists is not enough. Earlier versions of this guide said the safeguard was a human checking every citation against the primary source, and that sentence, followed literally, still gets lawyers sanctioned. Look at what the 2026 record actually punishes. Couvrette involved fifteen nonexistent cases and eight fabricated quotations. Doiban involved fifteen false citations and nine contrived quotations. Whiting involved fake citations and misrepresentations of the record.[258][251] Existence-checking catches the first failure in each pair and none of the second. Run four checks on every citation in anything you file:

  1. The case exists, verified in Westlaw, Lexis, or the court's own docket.
  2. Every quoted passage appears verbatim at the page cited.
  3. The case actually supports the proposition it is cited for.
  4. The case is still good law, checked in KeyCite or Shepard's.

Check the assigned judge's standing order before every filing. Many courts and individual judges have adopted orders on AI use, and they are not uniform. Some require a certification that AI-assisted content was verified, some require disclosing which tool was used, some require identifying the AI-assisted passages, and some prohibit AI use in submissions outright. They vary judge to judge inside a single district. This is not something to leave to a lawyer's memory, so make it a field in the docketing system alongside the page limits and the font rules.[282] A proposed Federal Rule of Evidence 707 would have applied a reliability gate to machine-generated evidence. On May 7, 2026, the Evidence Rules Committee declined to advance the proposal, substantially revised it, and held it for further study. It has no effective date and appears on no pending amendments list.[282]

Have a protocol for the fabricated cite you discover after filing. Model Rule 3.3(a)(1) addresses a lawyer's knowing false statement to a tribunal and the failure to correct a prior material false statement. Paragraph (b) applies when the lawyer knows that a person is engaging or has engaged in criminal or fraudulent conduct related to the proceeding. It does not govern every citation mistake.[83]

The response depends on the jurisdiction, posture, materiality, client duties, court rules, and insurance policy. Stop relying on the affected submission, preserve relevant records, verify the scope of the error, and obtain local ethics or malpractice advice on correction, client communication, court notice, and carrier notice. Treat that sequence as a risk protocol, not as a claim that one rule requires immediate notice to every recipient in every case.

4.8 Realistic adoption roadmap: 90 / 180 / 365 days

Before any of what follows, audit your outside counsel guidelines. Some clients' guidelines may restrict AI use, require consent or vendor disclosure, or change what may be billed. I did not locate a defensible prevalence survey, so this version makes no claim about how common those provisions are.[283] Read the actual guidelines for your largest clients, decide what requires notification or consent, and do that before signing the vendor agreement.

A staged rollout beats a firmwide switch, and the right shape is pilot, then scale, then build. The pattern below is a generic adoption sequence adapted to a firm's risk posture, not a sourced law-firm roadmap.

In the first 90 days, get the governance and a narrow pilot in place. Sign an enterprise agreement with zero-retention and no-training terms, stand up an AI use policy grounded in Opinion 512, and pilot one well-bounded workflow, document summarization or first-pass research, with a single practice group. Measure real time saved, not the feeling of speed.

By 180 days, expand what worked and formalize review. Widen the pilot to more groups, build a shared prompt library and a context base of the firm's templates and standards, and make the human-verification step a documented part of the workflow rather than an informal habit.

By 365 days, decide on the build-or-buy question with evidence in hand. By now the firm knows which workflows deliver value and at what cost, which is the only sound basis for deciding whether to wrap a model, invest more heavily in a vendor, or, for the largest firms, build. Adoption discipline matters throughout: tools that sit unused deliver nothing, so usage is the metric that determines return.

4.9 The Finnegan / Davis Polk / Desmarais comparison

These three firms illustrate three distinct postures, and the contrast is instructive precisely because not all of them have disclosed much.

Finnegan, an intellectual-property firm, sits at the most documented end. In September 2025 it launched a formal client-facing practice, AI + Finnegan, with four teams spanning AI and patents, copyright, privacy, and trade secrets, and named partner leads.[204] On the internal side, the firm's own disclosures describe an AI tech-tools committee of partners, technologists, and risk staff, a knowledge-management and innovation team, and prompt-engineering training for its lawyers.[205] Those internal details come from Finnegan's own materials rather than outside reporting, and the firm has not named the specific tools it uses, so read the program as the firm presents it.

Davis Polk's public posture leans toward thought leadership and advisory work. Named partners write and speak on AI across the industry, and the firm runs a standing AI advisory practice, but it has not disclosed a flagship internal build.[78] That is a legitimate strategy: lead the client conversation, and adopt deliberately.

Desmarais LLP is the honest negative case. It is an IP-litigation boutique known for alternative fee arrangements, and the available research found insufficient public information about any specific generative-AI deployment. Rather than infer a strategy from silence, the accurate statement is that its posture is not publicly documented.

The comparison's real lesson is about evidence. Two of these firms have visible programs and one does not, and a firm evaluating its own path should weigh what peers actually disclose against what gets asserted in vendor marketing.

The practical future of legal AI is not a blank chat box with a better model behind it. It is packaged workflow assets: matter-intake checklists, citation verification harnesses, privilege-review procedures, redaction pipelines, research rubrics, and approval logs. That is the difference between a lawyer experimenting and a firm operating.

The June scrape reinforced this. The strongest legal signals were not "ask the model to be a lawyer" prompts. They were repositories, templates, and agentic systems that encode how legal work should move from source material to reviewed artifact.[215] A firm should treat these assets like practice infrastructure: version them, review them, assign owners, and test them on known examples before using them on live matters.

When a vendor pitches workflow assets, look for the parts that survive contact with a malpractice question. A short checklist:

If a vendor cannot show these, the asset is a demo, not infrastructure.

The clearest example of packaged workflow assets arrived on May 12, 2026, when Anthropic released Claude for Legal: an open-source suite of practice-area plugins covering commercial, employment, privacy, corporate, and AI governance work, a large set of named workflow agents, and more than twenty connectors into the systems firms already run, including iManage, NetDocuments, Ironclad, DocuSign, LexisNexis, Thomson Reuters CoCounsel, Relativity, and Everlaw.[248] Freshfields, Quinn Emanuel, and Holland & Knight were named as using Claude on live matters at launch.[248] It is available to paid Claude customers rather than sold as a separate tier, and enterprise administrators control which pieces are enabled.[248]

Three things about this matter more than the product itself.

It resets the price of the workflow layer. When practice-area scaffolding ships open-source from the model provider, a vendor charging separately for the same scaffolding has to justify the premium on something else: proprietary data, verified retrieval, insurance, support, or an audit trail you can hand a regulator. That is a useful question to put to any legal-AI vendor in a procurement conversation now.

It moves the risk from drafting to configuration. Plugins are starting points, not firm policy. A plugin encodes somebody's general view of how employment counsel work should proceed, and your firm's playbook is not that view. Treat a plugin the way §4.10 says to treat any workflow asset: version it, assign it an owner, adapt it to firm precedent, and test it on matters where you already know the right answer.

It makes connector permissions the live control. Several of these connectors write as well as read, into e-signature and document management systems that hold the operative versions of client documents. Before enabling any of them, confirm the permission scope, confirm the zero-retention and no-training terms the confidentiality analysis in §4.6 requires, and put a named human approver in front of any action that sends, files, or executes. A read-only connector is a research tool. A write-enabled connector is an agent acting in your client's name.

4.11 Litigation workflows worth running now

Everything above is about governance. This section is about the work. Six litigation workflows moved from demo to production in 2026, and each one has a named tool and a specific failure mode a litigator has to own.

Brief drafting with a citation-verification harness. Thomson Reuters launched the next generation of CoCounsel Legal on August 20, 2026, including a Westlaw brief builder and a verification pass that checks cited propositions against the underlying source.[255] This is the first widely available product that does what §4.10's checklist demands rather than describing it. Do not mistake grounding for immunity. The Stanford RegLab study remains the only preregistered public benchmark of citation-grounded legal tools, and it found hallucination rates in the 17% to 33% range.[255] No independent 2026 re-test has published, so the vendor's accuracy claims are unaudited.

E-discovery review, now bundled rather than premium. Relativity made its generative review and privilege-detection features standard in RelativityOne, Everlaw made single-use AI tasks free, and DISCO shipped an agentic tool that plans and runs multi-step fact investigations.[256] The economics of first-pass review changed as a result, which is the single largest cost line in most litigation budgets. The control does not change: an agent that autonomously searches, filters, and ranks still needs QC sampling against a validated seed set before anyone certifies a production as complete, and a privilege call remains a lawyer's call.

Deposition preparation and transcript analysis. CoCounsel, DISCO, and Everlaw all now generate deposition outlines and witness-testimony summaries.[255][256] The failure mode here is quiet and dangerous. A summary that omits an impeaching contradiction or softens an admission is worse than no summary, because it produces confidence rather than doubt. Spot-check every summary against the transcript before it shapes an examination.

Chronology and fact development. Agentic tools now extract dated facts from a document set and assemble a source-cited chronology.[256] Errors cluster where you would expect: a fact on the wrong date, or attributed to the wrong document, because the model misread a Bates stamp or a poorly scanned exhibit. Verify each entry against the page before a chronology reaches a mediation brief.

Judge and opposing-counsel analytics inside general chat tools. Trellis connected its state trial court dataset to Claude in May 2026 and to ChatGPT in July 2026, which moves analytics that used to live in a specialist platform into the tool a lawyer already has open.[257] Treat the outputs as base rates rather than forecasts. Using a grant-rate statistic to set a client's expectations or settlement authority, without saying out loud how thin the matching sample is, is a Rule 1.4 problem waiting to happen.

Medical record and damages review. First-pass extraction of medical records into litigation-ready chronologies is now a product category rather than a service.[256] The asymmetry is brutal: an AI that misses a pre-existing condition hands the gap to opposing counsel, who will find it. Clinical human review of flagged records is not optional, and vendor claims of 40% to 60% cost reduction are sales figures, not audited benchmarks.

One pattern runs through all six. In every case the tool compresses the first pass and leaves the judgment, and in every case the failure mode is a confident omission rather than an obvious error. Omissions do not announce themselves the way a fabricated citation does. Build the sampling step into the workflow, because you will not notice you needed it.


Part 5: Building with code

This is the part I live in. If you write software, or you want to, the AI tools have moved from autocomplete to something closer to a junior engineer who never gets tired. The catch is that the people getting the most out of them are not the ones with the cleverest prompts. They are the ones who set up the workspace carefully before they ask for anything. Everything below is about that setup.

5.1 The philosophy: plan first, scaffold deliberately

The single biggest predictor of a good result is how much context the model has before it starts, so I spend more time on setup than on the prompt itself. Anthropic calls this context engineering, and the framing is useful: the model's context window is a budget, and your job is to spend it on the few things that actually steer the output toward the right answer.[115] A model with the right files in front of it and a clear goal beats a smarter model working blind.

Plan before you generate. For anything past a one-file change, I write down what I want, let the model draft an approach, and only then turn it loose on code. The cost of a wrong plan is a few minutes of reading. The cost of a wrong implementation is an afternoon of unwinding it.

This is not a fringe practice anymore. Anthropic reports that as of May 2026, more than 80% of the code it merges into its own codebase was authored by Claude.[230] That number only holds because the work is scaffolded and reviewed, not because the model is turned loose unsupervised. The scaffolding is the point.

Scaffold deliberately means the project has a shape the model can read. Clear file names, a short description of what each part does, and a written goal. The model is good at filling in a well-defined frame and bad at inventing the frame from a vague request. Give it the frame.

In production, the prompt is only one surface. The schema, tool description, validator, memory layer, and eval set are also part of the prompt because they all shape what the model can do. Use the prompt for business meaning, audience, evidence rules, and edge cases. Use schemas for machine-readable shape. Use validators and evals to catch the parts language alone cannot reliably enforce.

5.2 Starter pack: design.md, identity.md, PRD, AGENTS.md, CLAUDE.md

A handful of plain markdown files in the repo root will do more for output quality than any prompt trick, because they turn one-off instructions into standing context the model reads every time. Andrej Karpathy popularized the idea of keeping an LLM-readable wiki for a project, a living set of notes written for the model rather than for a human team.[56][128] Here is the set I keep.

This is the coding version of the same Projects rule from Part 3. If the agent has file access, put context in files and direct it to those files. Do not paste the same design notes, screenshots of requirements, or long source files into the chat every time. The repo is the project memory. The chat is the workbench.

design.md is the architecture in prose. What the system does, the major pieces, how they fit, and the decisions you have already made so the model does not relitigate them.

identity.md is the voice and the non-negotiables. For a writing project it is the style guide. For a product it is the tone, the naming conventions, and the things you never want the model to do.

A PRD (product requirements document) is the what-and-why of the feature you are building right now. Even a half-page keeps the model aimed at the real goal instead of a plausible adjacent one.

AGENTS.md is the cross-tool instruction file. It has become a rough standard that several coding agents read, so it is the right place for build commands, test commands, and conventions that any agent touching the repo should follow.

CLAUDE.md is the Claude Code equivalent, read automatically at the start of every session. Mine holds routing rules, tool preferences, and the mistakes I do not want repeated. These files override default behavior, so the more precisely you write them, the less you have to repeat yourself.

You do not need all five on day one. Start with CLAUDE.md or AGENTS.md and a design.md, and add the others when you notice yourself giving the same instruction twice.

As the project grows, add decisions.md for choices already made, open-questions.md for unresolved issues, sources.md for provenance, templates.md for reusable prompts or formats, and handoff.md for the current session state. Add files because repeated context needs a home, not because a checklist says every repo must have them.

5.3 Plan mode: when it beats just-do-it

Use plan mode whenever a wrong first attempt would be expensive to undo, and skip it when the task is small and reversible. Claude Code's plan mode has the model research and propose an approach without touching files, so you sign off before any code changes. The research bundles do not give a single decision rule here, so this is my own working heuristic.

Plan first when the change spans several files, introduces a new system, or touches anything you cannot easily roll back. The plan is cheap to read and cheap to correct, and catching a bad assumption there saves a full rewrite later.

Just do it when the task is one file, well-bounded, and easy to revert with git. Forcing a planning round on a two-line fix is its own kind of waste. The tell is reversibility: if git can undo it in one command and you would notice the mistake immediately, skip the ceremony.

5.4 Ralph loops and /goal: when iteration beats one-shot

For work where the right next step depends on what the last step produced, an iteration loop beats a single big prompt. The pattern that made this concrete is the Ralph loop, named and popularized by Geoffrey Huntley: you give the agent a goal and a stop condition, and it runs the same prompt against its own evolving output until the goal is met.[62] People have driven runs for many hours this way, with the agent fixing, testing, and re-fixing without a human in the seat for each turn.[61]

Anthropic later shipped an official ralph-wiggum plugin for Claude Code that formalizes the technique through a stop-hook, so the loop is now a supported pattern rather than a clever hack.[187] On the Codex side, the /goal command plays a similar role, letting you hand over an objective and let the tool iterate toward it.[58]

The danger with any unattended loop is that it runs forever and burns tokens with nothing to show. So every loop I run has four guardrails: a hard turn budget, a wall-clock deadline, a completion marker the loop can detect to know it is done, and a forward-progress check that halts the run if two turns in a row produce nothing new. Make that progress check read a number the model does not write: passing tests, items remaining, diff size. I tested the alternatives on synthetic loop histories, and plain repeated-output detection caught only one of four stall shapes, missing oscillation between two states and the expensive case where the agent produces different plausible text every turn while nothing improves. A July 2026 static analysis backs the worry with found defects rather than a rate: scanning 6,549 agent repositories, it confirmed 68 unbounded-loop failures across 47 projects, each one a repeated path that no effective bound covered.[296] It analyzed source code, not live runs, so it shows the gap exists, not how often it bites. Without these guardrails, an agent loop can quietly spend an enormous amount of money making no progress. With them, iteration is the best tool for fuzzy, long-horizon work.

One boundary to be honest about: these guardrails bound waste, not blast radius. Nothing in a turn budget stops a loop from deleting the wrong thing well inside its budget. Action-level safety comes from §5.8: scoped credentials, sandboxes, separation of production from everything an agent touches, and approval gates on anything destructive or externally visible.

The honest test for whether to loop at all: if the task could plausibly succeed on a scheduled runner with no model in the loop, write a plain script instead. Save the agent loop for work where the success criterion is not yet mechanical.

5.5 Context management: compact, handoff, artifacts, and vaults

Context is the scarce resource, and managing it well matters more than picking a slightly better model. Models degrade as their context fills, an effect Anthropic and others have documented as context rot: the same model gives worse answers late in a long, cluttered session than it does early in a clean one.[126][115] So the skill is keeping the window full of the right things and empty of everything else. Appendix G distills this section into a checklist you can run before any long chat or agent run.

Three moves cover most situations. Compact when you need to keep going in the same session but the window is heavy. A targeted compaction that preserves the slice you care about and crushes the rest is better than a blanket one. Handoff to a fresh session when the current task is self-contained and the accumulated context would only confuse it: write a short brief, commit any pending work first, and start clean. Fresh window between unrelated tasks, always. Shared context across two different jobs degrades both.

The handoff is the habit that makes fresh windows painless. Before moving to a new chat, ask the model to write a self-contained brief: goal, current state, files or project context to read first, decisions already made, what not to redo, open questions, risks, and the next concrete action. Paste that into the new window instead of dragging the whole old transcript with you.

Compaction itself is lossier than it feels, and now there is a measurement. One practitioner ran a decay probe on a small model (Claude Haiku 4.5 through the Claude Code command line): seed ten synthetic facts, delete their source file, force repeated compactions. Zero of ten facts survived the first boundary, and 106 of 108 later summaries carried none. The two exceptions prove the rule. The agent's own grep accidentally re-read the facts from a file on disk, they returned for two summaries, and one re-summarization later they were gone again. The summarizer even reported it had kept all ten while carrying zero.[301] One author, one model, one CLI version, so hold the exact counts loosely, but the split is the lesson: a fact that must survive belongs in a file the session re-reads, or in a hook that re-injects it at every boundary, never in the conversation alone.

The general principle behind these moves is the same one from 5.1. The window is a budget. Long-running agents that succeed are the ones that keep retrieving the right context at the right moment rather than carrying everything forward.[116][118]

Persistent artifacts are the fourth move. A knowledge vault gives the agent a place to put durable facts, decisions, open questions, and source notes so the next session does not have to rediscover them.[209] For complex reviews, an HTML artifact can outperform a markdown transcript because it can hold a plan, diff, tables, screenshots, diagrams, and reviewer controls in one navigable surface.[207] The rule is simple: keep conversation light, put durable work in files, and make the next agent read the file instead of the whole chat. One warning before you build the vault: files an agent re-reads as instructions are also an attack surface. Section 5.8 covers memory poisoning and the review discipline that answers it.

The budget principle has one measured exception at the small end. For material an agent needs on every turn, a small always-loaded index beats knowledge gated behind a tool the agent must decide to invoke: on one vendor's eval of deliberately post-cutoff framework APIs, an 8KB always-loaded index passed every case where the gated version plateaued at 79%, because the agent often never decided to look.[302] At library scale the economics flip toward single-level retrieval, and adding a deeper routing layer can actively hurt: a July 2026 controlled study measured one harness's multiple-choice accuracy collapsing from 0.91 to 0.64 under an extra hierarchy, while at twenty-book scale flat indexing reached nearly double the open-ended accuracy at half the cost.[303] Small and always-needed loads up front. Big and occasional gets one flat index, retrieved when needed. Test before adding anything deeper.

Skill graphs need one more routing rule. If the knowledge is shared across projects, keep it in a standalone knowledge base and point agents to it. If the knowledge belongs to one product, keep it in that repo's docs or a repo-local knowledge graph. If you put a graph under .claude/skills/, use a small index or trampoline file and keep deep knowledge nodes out of the skill-description budget. Skills should trigger procedures. Knowledge files should hold facts. When those two blur, agents start loading the wrong thing at the wrong time.

The routing rule is:

Situation Best move
Same task, same context, window getting heavy Compact with a focus
New task, old context mostly irrelevant Fresh window
Same project, different phase Handoff prompt into a fresh window
Long research or many file reads Separate worker or subtask if available
Output needs to persist Save an artifact or file
Chat has gone wrong Fresh window with corrected handoff
Recurring work Project files, not repeated uploads

I also like a lightweight tripwire: after roughly 25 prompts, or before starting a heavy new subtask, ask whether the next move belongs in the current window at all. If the work is mechanical, delegate it. If it will produce noisy search output, send it to a separate worker. If the current window is full of unrelated history, write a handoff and start fresh. If none of those apply, stay put.

5.6 Sub-agents and tier sizing

Delegate independent subtasks to sub-agents, and match the model tier to the job rather than reaching for the biggest model every time. A sub-agent runs a scoped task in its own context window and returns just the result, which keeps the noisy middle of a search or a bulk edit out of your main session.[117] When several subtasks are independent, you can run them at once.

Tier sizing is the discipline that keeps this affordable. The rule I use is to pick the cheapest model that can do the subtask well:

Two guardrails keep the pattern from sprawling. Cap the spawn depth so a sub-agent can spawn at most one further tier, and do not escalate to a bigger model without a concrete reason. The parent session owns the final answer and stitches the pieces together. Most of the time the right move is several Sonnet agents in parallel, not one Opus agent doing everything.

Tool support varies. Some coding tools have native sub-agents. Others need you to launch separate CLI jobs, background workers, or fresh sessions. The principle is the same either way: isolate noisy work, give each worker a narrow brief, and bring back only the result.

The orchestration step is now automated in at least one tool. Anthropic announced dynamic workflows for Claude Code on May 28, 2026, and they are now generally available: instead of you deciding how to split the work, the model writes an orchestration script and runs tens to hundreds of parallel sub-agents in one session, including agents whose job is to refute what the other agents found, iterating until the answers converge.[245] I used this pattern to pressure-test this version of the guide, and it caught a pricing error I had introduced from a secondary source.

Two cautions before you reach for it. Anthropic's own documentation warns that dynamic workflows consume substantially more tokens than an ordinary session.[245] And fan-out is the wrong shape for work that is genuinely sequential, where each step depends on the last. Use it when the work splits into independent lanes or when you want several independent readings of the same question. Use a single session when the work is a chain.

Skills and sub-agents solve different problems. A skill is reusable know-how for a class of work: "prepare a diligence memo," "audit accessibility," or "run the team release checklist." A sub-agent is a temporary worker for one scoped task. Use skills when the same workflow will recur. Use sub-agents when the work can split into independent lanes. Use both when a sub-agent needs a known procedure.

5.6.1 Mixed-model and multi-agent review

Use multiple models when the task has judgment risk, hidden assumptions, or several plausible answers. The pattern is generate separately, compare structurally, then synthesize. Ask Model A and Model B for independent answers, keep their outputs separate, and run a comparison pass that extracts agreement, conflicts, missing evidence, and unique ideas. OpenRouter's Fusion description is the productized version of this workflow: fan out to several models, then use a judge to identify consensus and contradictions.[206] One caution on that product shape: the consensus half is the half the measurements below undercut. Use the fan-out for the contradictions it surfaces, not for the comfort of agreement.

Do not use model agreement as proof. A study of over 350 models found that on one leaderboard dataset they agreed about 60% of the time when both were wrong, and that error correlation rises as models get more capable, across vendors and architectures.[291] A separate 2026 study across five models and four truthfulness datasets found no aggregation trick, majority voting included, reliably beat asking one model once, even at up to 25 times the inference cost.[292] An audit of three-model panels on a standard benchmark found voting beat the panel's own strongest member less than 10% of the time.[293] Agreement is not a check. Disagreement is a usable flag, though: in one structured clinical-extraction study, the lowest-agreement 6.5% of items held roughly 80% of all annotation errors, so routing just that slice to human review would have caught most mistakes at a fraction of the effort. One site, one language, one narrow task, and the authors say so.[294] The value of mixed-model review is that it exposes uncertainty faster than one chat does. It is most useful for research synthesis, legal issue spotting, architecture tradeoffs, strategy memos, and any task where "one fluent answer" is too smooth to trust.

For code and data work, the strongest version is writer and reviewer separation. Have one agent write the patch, migration, or analysis. Have a second agent from another model family review only the diff, tests, contracts, and likely edge cases. The reviewer should not rewrite the work by default. Its job is to say: ship, fix these specific issues, or stop because the risk is higher than the change is worth.

The mechanism behind the separation matters. A second look inside the same session shares the writer's context and catches little, the self-correction limit measured back at ICLR 2024: without external feedback, models fail to find their own reasoning errors and sometimes get worse on a second pass.[192] So separate the review context, prefer another model family, and ground the verdict in something external, tests, retrieval, a compiler. And treat the model reviewer as a biased instrument, not a neutral one: in a peer-review simulation, model reviewers accepted model-written papers at 78% against 49% for human-written papers on identical topics.[295]

5.7 Where to host: AWS vs Cloudflare vs local LLM

Match the host to the workload: a big cloud for heavy or compliance-bound systems, an edge platform for lightweight global apps, and a local model when privacy or cost rules out sending data anywhere. The cloud comparison below is general guidance rather than a benchmarked recommendation.

AWS and the other large clouds make sense when you need deep service integration, heavy compute, or specific compliance certifications, and you are willing to manage more infrastructure to get them. Cloudflare and similar edge platforms suit lightweight, globally distributed apps where you want low latency and minimal ops, with the tradeoff of a more constrained runtime. For most small projects, the edge option is the faster path to something live.

Local LLMs are the interesting third option, and the tooling is now good enough for real use. Ollama is the easiest on-ramp for running an open model on your own machine, LM Studio gives you a friendly interface, and vLLM is the choice when you need serving throughput: on Red Hat's benchmarks it served roughly 793 tokens per second at peak against Ollama's 41, an order-of-magnitude gap that only matters once you are serving many users at once.[152][153] The open models worth running locally are mostly Meta's Llama family and a handful of others, which are capable enough for many tasks without an API call leaving your hardware.[38] You host locally when the data cannot leave, when you want to eliminate per-token cost, or when you simply want to own the whole stack. The price is that you manage it, and the best open models still trail the frontier hosted ones.

5.8 Security: secrets, .gitignore, MCP risk surface, prompt injection

Treat anything an AI agent can read or run as part of your attack surface, because it is. The basics first: keep secrets out of the repo, put .env and credential files in .gitignore, and never paste live keys into a prompt. Stage specific files by name rather than committing everything blindly, so a stray secret does not ride along.

MCP (Model Context Protocol) is the connector standard that lets agents reach external tools, and it widens the risk surface in ways worth understanding. Security researchers have catalogued real problems: authentication that the protocol treats as optional, over-broad tool permissions, token handling that exposes credentials, tool-description poisoning, and a remote-code-execution pattern in a widely used MCP SDK that OX Security disclosed in April 2026. The researchers estimated more than 7,000 public deployments in the affected supply chain, but execution depended on downstream software accepting attacker-controlled STDIO configuration. That estimate is not a count of proven compromises. Anthropic characterized the underlying behavior as expected rather than shipping a protocol-level patch, which puts the burden on you to lock down each deployment.[100][137][135] The NSA's May 2026 security guidance flags that an MCP session can carry broad ambient authority, where a connection authenticated once keeps acting with that access across a whole session, so harden scope and permissions at deployment time rather than assuming the default is safe.[111] The practical rule is least privilege: give a connector only the access it needs, and review what each one can actually do before you trust it.

Two things changed over the summer, one good and one alarming. The good one: the MCP specification revision published July 28, 2026 rewrites the protocol as stateless request and response rather than long-lived bidirectional sessions, hardens OAuth with issuer validation and audience-bound tokens, and deprecates dynamic client registration in favor of client ID metadata documents.[259] That revision targets precisely the ambient-authority problem the NSA flagged, so if you deployed MCP before August 2026, the upgrade is a security fix and not a feature release.

The alarming one: the first large-scale dynamic audit of internet-facing MCP servers, published July 31, 2026, found more than 21,000 discoverable servers, and of those it audited dynamically, 91.8% had no OAuth at all and 687 tool instances exposed raw shell access.[260] Those are not misconfigured hobby projects on the margin. That is the modal state of publicly reachable MCP infrastructure. Before you connect an agent to a third-party MCP server, assume it is unauthenticated unless you have checked.

Prompt injection is the one to internalize. A malicious instruction hidden in a web page, a document, or a tool result can hijack an agent that reads it, turning your own assistant against you. It sits at the top of the OWASP list of large-language-model risks for good reason.[114] You cannot fully prevent it yet, so the defense is structural: do not give an agent both untrusted input and broad, irreversible permissions at the same time, and keep a human in the loop for anything destructive or externally visible. If a tool result looks like it is trying to give you instructions, treat that as the alarm it is.

That structural rule has a gap, and researchers found it in July 2026. Agent data injection plants fake delimiter and punctuation characters inside structured data fields rather than planting instructions in prose, which corrupts what an agent trusts as legitimate metadata rather than what it reads as text.[261] The disclosed technique worked against Claude Code, Claude in Chrome, OpenAI Codex, Gemini CLI, and Google Antigravity in their default configurations.[261] Watching for instruction-shaped text does not catch it, because nothing in the payload looks like an instruction. The mitigation is a different category: sanitize the boundary between trusted and untrusted data, and do not let a data field silently become a control character.

The most durable variant is memory poisoning, and it targets exactly the file-based memory this guide recommends. Every file an agent auto-loads as instructions, memory files, CLAUDE.md-class project files, knowledge wikis, skills, is a persistence layer an attacker can write to once and profit from every session afterward. In April 2026, Cisco researchers demonstrated a malicious npm package appending attacker instructions to a Claude Code memory file during install. The poisoned lines loaded at every session start, and the compromised agent recommended insecure practices, committing API keys to source among them, persistently and with nothing visible to the user. Anthropic changed how memories load in response, in Claude Code v2.1.50.[307] Two related Claude Code vulnerabilities, both since patched, ran hook commands from project files the moment a workspace was trusted, with none of the per-command confirmation a user would expect, and redirected API traffic, credentials included, through an environment-variable override.[308] The defense is to treat memory and knowledge files as code: keep them in version control, review their diffs like any other change, and audit anything that can write to them, package installs and hooks first.

Vendors are starting to ship this structural defense as a product setting. OpenAI's Lockdown Mode, rolling out from June 4, 2026, disables web browsing, image generation, Deep Research, Agent Mode, connectors, and downloads in one switch, which removes most of the channels an injected instruction needs to do damage, and its Elevated Risk labels flag sessions where untrusted content raises the exposure.[228] Anthropic shipped the structural counterpart on August 5, 2026: inference hooks, in beta, hold every governed prompt across Claude, Cowork, and Claude Code for an allow or deny verdict from the organization's own security server before inference proceeds, and log denials to a compliance activity feed.[262] For a regulated organization, that is the difference between a policy document and an enforced control, and it is the kind of thing to ask about in procurement.

The lesson for your own builds is the same whether you buy it or build it: the fewer capabilities an untrusted-input session can reach, the smaller the blast radius. Anthropic's own containment writeup publishes the reason the model layer cannot stand alone: on an external red-teaming benchmark, for one model version, injection resistance of roughly 0.1% attacker success on single attempts rises to about 5 to 6% under a hundred adaptive attempts.[309] So deterministic boundaries back the probabilistic one: an OS sandbox that denies network by default and confines writes to the workspace, plus a separate approval gate for actions beyond it. Those figures are the vendor's own. The architecture lesson stands without them: layer the model, the sandbox, and the treatment of external content as adversarial, so no single miss is a full compromise.

5.9 Agent primitives: browser, tools, skills, sandboxes, evals

A lot of useful capability is available as free or open APIs, but the deeper shift is that agent systems now have reusable primitives. Models are one primitive. Tool calls, browser sessions, skills, sandboxes, eval manifests, and knowledge stores are the others.

Browser automation is the clearest example. A browser skill can package selectors, XHR/API knowledge, login/session expectations, and a repeatable workflow so an agent does not rediscover the site every time.[212] Skills do the same for local work: a scoped instruction bundle tells the agent when to trigger, what context to read, and how to perform a narrow class of work. Anthropic's official skill docs make the storage pattern concrete: a skill can carry instructions, scripts, and resources, and good skills stay concise, structured, and tested against real usage.[237] Sandboxes keep code execution and untrusted inputs away from secrets. Eval manifests make the result repeatable by recording the prompt, model, tools, inputs, and scoring rubric.[210]

Tool descriptions deserve the same care as user prompts. A good tool description says when to use the tool, when not to use it, what inputs are required, what precondition must be true, and what the model should do with the result. Sloppy tool descriptions are a quiet source of bad agent behavior because they teach the model the wrong affordances.

Some of these primitives now ship built into the model APIs themselves. Google's Gemini API exposes a URL Context tool that fetches and grounds on the pages you name, and OpenAI's Responses API ships built-in web_search and computer-use tools, so retrieval and browsing become a flag rather than a separate service.[229] Google has also moved a full agent harness into the API: Managed Agents can run in isolated Linux environments, execute code, use tools, and be defined with AGENTS.md and SKILL.md files.[236] Reach for the built-in primitive before you wire up your own when the provider already offers it.

On the model side, the open-weight families, mostly Meta's Llama models, can be run yourself or accessed through hosted providers, which gives you a fallback that does not depend on a single vendor.[38][39] Beyond models, the pattern that pays off is connecting an agent to the specific data sources your work depends on through MCP, so the model is reasoning over your real information rather than its training data. The pricing and exact capabilities of individual public APIs move too fast to pin down here, so check the current terms before you build on one.

5.10 Operator patterns from the field

The best workflow ideas often appear in repositories and practitioner forums before they reach official documentation. Use those places as an intake queue, not as an authority. The companion Operator Field Notes keeps the fast-moving layer separate from the durable guide and grades every pattern as reproduced, corroborated, promising, anecdotal, or retired.

Six patterns have earned a place here: five on independent corroboration, and the prompt compiler on its mechanism, with its efficacy still an open question.

Compile the prompt before execution. Section 2.8 gives the everyday version. For serious work, let a planning model interview you and write a standalone prompt containing the objective, inputs, constraints, non-goals, output format, acceptance tests, verification steps, and stop conditions. Review that file, then give it to a fresh executor. A separate reviewer gets both the prompt and the result.[287] Label the evidence honestly: no published study measures this pattern against direct prompting, and heavy structure can hurt the newest models.[310] In my own small test the compiler surfaced constraints I had missed and still omitted the acceptance tests and stop conditions, so the review of the compiled prompt is the stage that earns its keep.

Run loops with state and backpressure. A reliable Ralph loop selects one bounded backlog item, starts a fresh session, loads the prompt and learned guardrails, makes one change, runs objective checks, records state, and repeats. The core mechanics, fresh context per task and small reviewed changes, are documented across independent repositories,[284] and the most complete implementation separates prd.json, prompt.md, and guardrails.md, then halts on repeated output or stalled progress.[285] Section 5.4 covers why the progress check must read an external metric.

Let a referee declare done, never the worker. Completion needs a machine-checkable convergence type. One practitioner essay names four: a named test scope passes, an iteration produces zero diff, a queue reaches zero, or a separate model scores against a rubric, the weakest of the four.[297] Two clarifications keep those four honest. Zero diff counts as convergence only when a passing acceptance check comes with it, because zero diff with a flat or failing check is the stall §5.4 halts on. And judge-defined ranks last only when the judge scores prose against a bare rubric. Section 5.6.1's reviewer earns more trust because its verdict is grounded in diffs and tests. The process that declares done-ness must not be the process that did the work. One practitioner puts it plainly: the model that writes the code is not the model that decides whether the code is done.[298] A community writeup of Codex's goal state machine describes the same split shipped: the model can create and complete a goal, but pause and budget transitions stay system-controlled.[299] For unattended loops, make the contract itself checkable: a public JSON Schema requires every loop to declare its verification receipts, budget, escalation, and exit conditions before it runs, and its runnable example logs each outcome as an append-only receipt with an evidence digest.[300] The schema checks shape, not sense, so a human still reads the contract once.

Make knowledge lintable, and know what lint cannot catch. The functional triad from §1.4.1 combines immutable raw sources, a model-written wiki, and a human-governed schema. Ingest, query, and lint are equally important. The linter catches unsupported pages, stale claims, broken links, and contradictions before the wiki quietly becomes a new source of confident error.[286] The linting ecosystem has matured: independent implementations now converge on similar rules, and the best-verified one hashes each cited raw source into the note's frontmatter, so the linter flags drift between what a claim cites and what the source now says.[304] Know the boundary, though. In my own reproduction, structural rules caught missing provenance, stale dates, orphan sources, and an invented statistic, and a quietly strengthened verb passed every automated check. Structural lint is mechanical. Overclaim detection above that level needs a model or human comparison pass, which is what the claim ledger below is for. One practitioner taxonomy names six kinds of wiki drift, among the most serious being citation drift, where a page still cites a source but the claim no longer matches it.[305] My own addition to that list is the self-citation loop: once the wiki references its prior synthesis instead of raw sources, internal consistency stops being evidence of correctness. Treat a wiki that reads more smoothly every month as a thing to lint, not a thing to trust.

Turn disagreement into retrieval. When agents disagree, split the dispute into atomic claims and fetch the primary source. Do not vote. When they agree, still verify any claim whose cost of being wrong is high. Shared training data lets several models repeat the same mistake, and §5.6.1 now carries the measurements: correlated errors rise with capability, aggregation fails to beat a single sample, and low agreement concentrates real errors well enough to direct review effort.[291][294] This pass supplied its own worked example. Two research agents returned opposite signs for the same statistic, and only re-quoting the primary source settled it.

Keep a claim-to-evidence ledger. For every consequential factual claim, record the exact supporting passage, source, retrieval date, scope of support, and unresolved contradictions. Citation resolution proves that a source exists. The ledger tests whether it supports the sentence. Recent systems research uses claim-evidence graphs and structured provenance interfaces for the same reason,[289][290] and a third independent group has now measured the approach: an evidence-ledger agent that labels each claim-evidence pair and routes unsupported claims back to the author during drafting correctly labeled 67.6% of pairs on a 2,335-row blind benchmark, against 38.3% for the best non-agent baseline.[306]

The remaining field patterns are operating disciplines rather than tricks: convert repeated corrections into tests or guardrails, give unattended work separate outcomes for verified completion and unverified completion, and increase autonomy through shadow mode with promotion gates written down before the shadow period starts. Compare tools by cost per accepted artifact rather than tokens or seats, with the label read out loud: it is an emerging discipline with one named proponent, not an industry standard.[311] The first practitioner benchmark supports the mechanism, one fixed feature across 23 model-and-harness combinations cost between $2.59 and $34.82 per merged result, and the cheapest model by token price was not the cheapest by merge. The same benchmark's review gate approved code failing up to 26 of 38 unit tests, so define "accepted" with hard checks before the metric means anything.[312] The companion holds the experiments and failure modes so this section can stay short.

The throughline of this whole part: the payoff is in the scaffolding, not the sentence you type. Set up the context, pick the right tier, guard the loops, and the tools will carry more than you would expect.


Part 6: Reference matrices

This part is the quick-reference. Three tables you can scan without reading the prose around them, each with an as-of date because all of it moves. When a table and your own test disagree, trust your test. Benchmarks age in weeks.

As of August 31, 2026. Pick by the job, not the headline ranking, because the leads are narrow and they rotate.[26] Prices are per million tokens, input then output. Every row here matches the model list in §1.2. If you find one that does not, §1.2 is the one I updated first.

Task First choice Why Price (in / out)
Coding, long agent runs Claude Opus 5 Near Fable 5 on coding benchmarks at half the cost, effort setting for tuning spend[239] $5 / $25[240]
Top-capability writing, reasoning Claude Fable 5 Anthropic's most capable widely released model[223] $10 / $50[240]
Broad knowledge work GPT-5.6 Sol Flagship of OpenAI's current family, GPT-5.5 held the top GDPval result[8][241] $4 / $20 promo at least through Nov. 21, later rate unannounced[241]
Everyday volume work GPT-5.6 Terra or Claude Sonnet 5 Mid-tier quality at roughly half the flagship price[240][241] $2 / $12 or $2 / $10
Multimodal (video, audio, PDFs together) Gemini 3.7 Flash One model reads every input type, cheapest frontier tier[242] $0.75 / $3.75, rises Jan 1, 2027[242]
High-volume / cost-sensitive GPT-5.6 Luna Cheapest frontier-family output by a wide margin[241] $0.20 / $1.20[241]
Long agent runs on a budget Grok 4.6 Built for long-running agents, but only a 500K context window[243] $2 / $6[243]
Live data, current events Grok 4.6 / Perplexity Grok for X data, Perplexity for citation-backed answers[243] varies
Legal work Claude for Legal plugins, CoCounsel Legal Practice-area scaffolding and source-grounded research, see §4.10.1[248][255] included / quote
Regulated domains (finance, healthcare) Claude for Financial Services / Healthcare Domain stacks with HIPAA-ready controls[233] quote / infra
Image generation Nano Banana 2 / ChatGPT Nano Banana for first image, ChatGPT for iterative edits[26] varies
Self-hosted / private Kimi K3 or Llama 4 Kimi K3 (July 27, 2026) is a far stronger open-weight option, under its own custom license, not Apache[264], Llama 4 remains the established family[38] infra cost only
Open-weight agent work Muse Glimmer 30B, Apache 2.0, genuinely unrestricted license[244] infra cost only

Two cautions on this table. Grok 4.6's 500K context window is half or less of every other frontier model listed, so it is the wrong pick when the job is reading one enormous document.[243] And Gemini's price is introductory: it doubles on January 1, 2027, so build the budget on $1.50 and $7.50.[242]

The reframe from Part 1 still holds: a frontier model can cost six times more per word than a cheaper rival for output that scores within a few points. Save the expensive model for the work that earns it.

6.2 Coding and agent-surface matrix

As of August 31, 2026. These are the coding agents and adjacent work surfaces: terminal agents that read your repo and run commands, AI-native editors, shared workspace agents, managed cloud agents, and browser-skill layers.

Tool Maker Shape Best for
Claude Code Anthropic Terminal agent, reads CLAUDE.md, uses skills, dynamic workflows Cross-file work, sub-agents, long autonomous runs, fan-out orchestration[1][237][245]
Codex (CLI) OpenAI Terminal agent with /goal loop Persistent goal-driven iteration[58][60]
Codex app OpenAI Desktop app supervising multiple agents Running and reviewing several agents at once[225][238]
ChatGPT workspace agents OpenAI Shared agents inside ChatGPT and Slack Scheduled team workflows with approvals and analytics[226][238]
Antigravity / Managed Agents Google Agent harness and cloud sandbox, remote control since Aug 20, 2026 API-driven agents with tools, code execution, and agent files[236][263]
Cursor Anysphere AI-native editor Editor-based workflow, inline edits[138]
Devin Desktop, formerly Windsurf Cognition AI-native editor plus cloud agents Pro $20, Teams $80/mo base plus $40/full seat[171]
GitHub Copilot Microsoft / GitHub Editor + CLI Microsoft and GitHub-centric shops[166]
Browser-skill surfaces Browse.sh and similar Reusable browser actions Repeatable web workflows and selector-heavy tasks[212]

Claude Cowork is a related but distinct product: a desktop agent that works across your files and documents rather than a repo-and-terminal coding tool.[217] It belongs on the consumer-agent rung of the ladder in §2.12. OpenAI's Codex app is the desktop counterpart to the Codex CLI: it lets you supervise several Codex agents at once rather than steering one terminal session, and it added Windows support in March 2026.[225]

One convergence is worth naming, because it changes where the work happens rather than how well it goes. Over the summer of 2026 all three major agent surfaces added remote or mobile session control: Claude Cowork reached web, iOS, and Android in July, OpenAI's Codex Remote went generally available across ChatGPT plans in late June, and Google's Antigravity added remote control on August 20.[263] The session keeps running on the machine that holds your files and credentials while you steer it from a phone. That is convenient, and it also means an agent with full local access is now reachable from a device that is easier to lose. Treat phone access to an agent session as equivalent to laptop access, and secure it the same way.

The pattern worth repeating from Part 5 is that agent surfaces are becoming file-shaped. Claude Code reads CLAUDE.md, Codex standardizes on AGENTS.md, Google's Managed Agents can use AGENTS.md and SKILL.md, and skills turn reusable procedures into files.[236][237] The scaffolding you write once steers every session.

6.3 Connector / MCP matrix (job to connector)

As of August 31, 2026. A connector, usually built on MCP (Model Context Protocol), is what lets an agent reach a real tool or data source instead of reasoning over its training data alone.[100] Match the job to the connection, and give each one only the access it needs.[111]

Job Connect to Notes
Search and draft over team docs Notion, Google Drive Read-scoped first, widen only when needed
Email triage and drafting Gmail, Outlook Keep send behind a human approval gate
Calendar scheduling Google Calendar, Outlook Low risk, good first pilot
CRM lookups and updates Salesforce, HubSpot Write access is higher risk, gate it
Code and issues GitHub Scope tokens to the specific repos
Team chat Slack Read for context, post behind approval
Browser workflows Browser session or browser skill Good for repeatable web tasks, avoid broad logged-in access
Managed code execution Cloud sandbox or Managed Agent Isolate untrusted code and keep environment state explicit[236]
Legal research and matter data CoCounsel Legal MCP, HighQ MCP Use permissioned, logged connectors for regulated workflows[231][232]
Legal practice workflows Claude for Legal connectors (iManage, NetDocuments, Ironclad, DocuSign, Relativity) Check each connector's write scope, and gate any send or file action behind a named approver[248]
Court and judge analytics Trellis MCP connector State trial court data inside Claude and ChatGPT, treat predictions as base rates, not forecasts[257]
Custom internal data Your own MCP server You own the permission model, least privilege

The security rule from Part 5 governs this whole table: never give one agent both untrusted input and broad, irreversible permissions at the same time, and keep a human in the loop for anything destructive or externally visible.[114]


Appendices

Appendix A: Glossary

Plain definitions for the terms used in this guide. One sentence each, in the sense this guide uses them.

Appendix B: Prompt and workflow patterns with examples

These are copy-paste templates for the techniques explained in Part 2. The full reasoning lives there. This is the quick card. Fill the bracketed slots.

For the fuller standalone version, see companions/prompt-engineering-mega-guide.md. For context management, handoff prompts, and project-file patterns, see companions/context-management-guide.md.

The universal prompt contract. Use when the task matters.

Help me [goal]. Context: [facts]. Audience: [who will read/use this]. Success criteria: [what good looks like]. Evidence rules: [what sources may be used]. Format: [desired shape]. If anything is uncertain, say so instead of guessing.

The meta-prompt. Have the model write the prompt, then run it.

Write me a prompt I can give an AI to [goal]. The situation is: [context]. I want [tone, length, and output format].

Show, do not tell (few-shot). Paste an example and ask for more in the same shape. The last example carries the most weight, so put your best one last.

Here are two status updates I liked: [example 1] [example 2]. Write this week's update in the same format and tone: [this week's facts].

Specify the output shape. Name the exact structure you want back.

Give me this as a table with three columns: risk, likelihood, and mitigation. No more than six rows.

Ask for graduated confidence. Get the model to mark what it is unsure about.

Answer, then mark each claim: confident, believe-but-uncertain, or insufficient information. Do not guess past what you know.

Provide sources, then ask for citations. Cuts fabricated references.

Using only the documents I pasted above, answer [question] and cite the specific document for each claim. Do not use outside sources.

Describe the method, not the credential. Beats "you are an expert in X."

Before suggesting a position, identify the other party's likely priorities and constraints, then propose a position that advances my goal while acknowledging theirs.

Generate, then self-critique. One or two passes only.

Now critique that draft. What is weak, incomplete, or possibly wrong? Then rewrite it to fix those issues.

Cross-check across two models. Run the same prompt through a second model. Agreement provides only weak evidence. Disagreement is a reliable flag to verify by hand.

Surface conflicts instead of smoothing them over. Use this when a source set or model panel may disagree.

Compare these answers/sources. Return four sections: consensus, contradictions, claims needing verification, and unique useful ideas. Do not synthesize until after the comparison.

Set a token or turn budget. Keeps agentic work from wandering.

Work toward [goal] for at most [N] turns or [N] tokens. Stop early if two consecutive turns make no new progress. At the end, report done, blocked, or needs human decision.

Set full agent guardrails. Use before an agent works for more than a few turns.

Goal: [objective]. Scope: [in bounds]. Out of scope: [do not touch]. Budget: [turn/token/time cap]. Checkpoints: after each major step, record what changed, what evidence supports it, and what remains. Stop if complete, blocked, two steps make no progress, or irreversible action is needed. Ask before any externally visible or destructive action.

Run writer/reviewer. Use when correctness matters more than speed.

Writer: implement [specific change] within [files/scope]. Run [checks]. Reviewer: read only the diff, contracts, and test output. Return: ship, fix specific issues, or stop. Do not rewrite the patch unless asked.

Close out agent branches. Use when agents create branches or result docs.

End the result doc with one closeout line: CLOSEOUT: APPLIED, WIP, SUPERSEDED-BY: [branch-or-commit], ABANDONED, or MERGED. No other states. The next session should know whether to keep, delete, or inspect the branch without reading the whole diff.

Turn research into schema. Use when sources should change a database, scorecard, rubric, or decision framework.

Extract research variables first: observable facts, metrics, dimensions, opinions, contradictions, source, and confidence. Group them into decision variables. Then list schema implications, scoring implications, missing data, and human decisions needed. Do not change schema from research alone.

Triage a long-running batch. Before running agents unattended, classify each task.

Task type Default handling
Docs, reports, isolated audits Safe to auto-apply if scoped
New ingest scripts or generated artifacts Auto-apply only after dry run and checks
SQL or migrations Manual review unless narrowly additive and verified
Edits to existing app behavior Manual review
Architecture, scoring, auth, billing, or production config Human decision first

Make an HTML artifact. Useful when the output needs navigation, tables, diagrams, diffs, screenshots, or reviewer controls.

Create a single HTML artifact that shows the plan, evidence table, open questions, and final recommendation. Make it reviewable without reading this chat transcript.

Compare models structurally. The everyday version of mixed-model review.

I will paste answers from multiple models. Compare them by agreement, conflict, missing evidence, risk, and best combined answer. Do not assume agreement means correctness.

Document a reusable team prompt. Use this before a prompt becomes shared operating infrastructure.

Create a prompt-library entry with: name, owner, use case, required inputs, model, prompt text, output format/schema, examples, known failure modes, review checklist, last-tested date, and version.

Describe a tool safely. Tool descriptions are prompts too.

Use this tool only to [specific purpose]. Do not use it for [nearby non-purpose]. Inputs must include [required fields]. Before calling, confirm [precondition]. After calling, summarize [result shape].

Start a fresh chat inside a Project. Use project files as durable context instead of re-uploading files.

Use the project files as durable context. Read [file] for [purpose], [file] for [purpose], and [file] for [purpose]. This is a fresh chat for [task]. If the project context is insufficient, ask what to read next.

Write a handoff prompt. Use whenever a new window or model takes over.

Write a self-contained handoff prompt for a fresh window. Include: goal, current state, files or project context to read first, decisions already made, what not to redo, open questions, risks, and the next concrete action. Keep it under 1,000 words.

Write a code handoff. Use before a fresh coding-agent session.

Write a handoff brief for a fresh coding-agent session. Include repo path, branch, current goal, what changed, what remains, files to read first with why, commands already run and results, known risks or failing tests, safety or don't-touch rules, and the exact continuation prompt. Do not include raw diffs or long logs.

Check citations and factual claims. Use before relying on research, legal, medical, or financial output.

Extract every citation, URL, case name, statute, statistic, date, price, and named source. Mark each as verified from provided source, needs manual verification, or likely unsupported. Do not invent missing citation details.

Recover from a bad answer. Use when the output is vague, too confident, stuck, or smooth but useless.

Re-review the answer for vague claims, overconfidence, missing evidence, unsupported assumptions, and unclear next steps. Return a table with issue, severity, evidence, and suggested fix. If the chat context is stale, produce a clean handoff prompt for a fresh window instead of continuing.

One caution carried from Part 2: telling a model "do not make things up" can backfire by making the unwanted behavior more salient. Say what you want instead.

Appendix C: Cost cheat-sheet

As of August 31, 2026. Prices churn every few months, so re-verify before you budget on them. See Part 1.3 for the reasoning behind the per-seat versus per-use split.

API token prices, per million tokens, input then output:

Model Input Output
Claude Fable 5 $10 $50[240]
Claude Opus 5 $5 $25[240]
GPT-5.6 Sol $4 $20 (promotional at least through Nov. 21, 2026, later rate unannounced)[241]
GPT-5.6 Terra $2 $12[241]
Claude Sonnet 5 $2 $10[240]
Grok 4.6 $2 $6[243]
Claude Haiku 4.5 $1 $5[240]
Gemini 3.7 Flash $0.75 $3.75 (introductory, doubles Jan 1, 2027)[242]
GPT-5.6 Luna $0.20 $1.20[241]

Discount mechanics are provider-specific. OpenAI distinguishes cache reads from cache writes and adds separate long-context input and output surcharges above 272,000 input tokens. Anthropic and Google use different cache, storage, Batch, and context rules. Price the exact request shape rather than multiplying the base rate by a universal discount.[239][240][241][242][243]

Consumer subscriptions:

Tool Free tier Paid
ChatGPT Yes, with limits Go $8/mo, Plus ~$20/mo, Pro $100 or $200/mo[17][253]
Claude Yes, with limits Pro $20/mo, Max $100 or $200/mo[184][253]
Gemini Yes, with limits US: AI Plus $9.99, AI Pro $19.99, AI Ultra $99.99 or $199.99[27][253]
Grok Yes, inside X SuperGrok, up to ~$300/mo Heavy tier[47]
Meta AI Free in Meta apps Free

Business per-seat:

Product Per-seat price
Microsoft 365 Copilot Business $18/user/mo, paid yearly, first-year promotion through Dec. 31, 2026[164]
Microsoft 365 Copilot Enterprise $30/user/mo, annual[164]
Microsoft Agent 365 $15/user/mo, agent governance control plane[277]
Cursor Pro $20/mo[138]
Devin Desktop, formerly Windsurf Pro $20, Teams $80/mo base plus $40/full seat[171]
ChatGPT Business Standard $20 annually or $25 monthly[185]
ChatGPT Enterprise Quote-only

The one number to remember: the spread between the most and least expensive output in that first table is more than forty to one, while the benchmark gap between the top and the middle is a few points. Match the model to the job.

Appendix D: Further reading

The primary sources worth reading in full, grouped by what they help with. Full citations are in the bibliography.

Model landscape, straight from the labs:

The two big shifts:

Building well:

Security:

For lawyers specifically:

Learning the tools:

Appendix E: Changelog

The canonical, maintained changelog is CHANGELOG.md in this directory, which records research imports and structural changes. The drafting history of the guide text itself:

Version Date Change
0.1.0 2026-06-01 Scaffold, front matter, audience selector
0.2.0 2026-06-02 Part 1 drafted (the AI landscape in 2026)
0.3.0 2026-06 Part 2 drafted (everyday user track)
0.4.0 2026-06 Part 3 drafted (regular company track)
0.5.0 2026-06 Part 4 drafted (large law firm track)
0.6.0 2026-06-08 Part 5 drafted (building with code)
0.7.0 2026-06-08 Part 6 drafted (reference matrices)
0.8.0 2026-06-08 Appendices A–F drafted, full draft complete
0.9.0 2026-06-08 Critique and integrity pass, disagreements resolved
0.9.1 2026-06-08 Citation and source-fidelity fixes
0.9.2 2026-06-15 DOCX render upgrade, source-fidelity and consistency fixes
0.9.3 2026-06-15 Klarna and vLLM throughput additions, version-cadence policy stated
0.9.4 2026-06-15 §3.1 rewritten vendor-agnostic with worked stacks, review cadence added, open questions resolved
0.9.5 2026-06-15 Resolved all outstanding verify markers: image and Siri claims corrected against research, editorial hedges kept in prose
0.10.0 2026-06-15 June scrape synthesis added: agent operating surfaces, mixed-model review, eval manifests, context artifacts, legal workflow assets, and agent primitives
0.10.1 2026-06-15 Prompt companion distilled into main guide: universal prompt contract, promptware library fields, production schema/tool guidance, and Appendix B templates
0.10.2 2026-06-15 Context-management companion added, main guide now emphasizes Projects/repo files as durable context, handoff prompts for fresh windows, and Appendix G checklist
0.10.3 2026-06-15 Firm-comparison §4.9 cited and tightened, price units fixed, unsupported Copilot dead-seat stat de-cited
0.10.4 2026-06-15 Audit remediation: Part 2 graduation ladder (§2.12) and intro added, malpractice framing in Part 4, §6.2 retitled Coding-tool matrix, price tables reconciled, citation-fidelity corrections (Anthropic opt-in, BCG/Dell'Acqua, Windsurf, Apple/Siri, Cowork). Pending: Bucket E content additions and companion-guide claim-level citations
0.11.0 2026-06-15 Model snapshot advanced to June 15 (Claude Fable 5 flagship, Gemini 3.5 Flash), product surfaces added (workspace agents, ChatGPT Apps, Codex app, Lockdown Mode, CoCounsel Legal + HighQ MCP, domain stacks, API research primitives, Anthropic >80%-code proof point), new §2.13 learning-the-tool track, companion guides given claim-level citations
0.11.1 2026-06-16 Agent-surface refresh anchored to official sources, skill graph, promptware, managed agents, and matrix updates
0.11.2 2026-06-16 Cross-repo workflow patterns added: writer/reviewer, context tripwire, branch closeout, long-running batch triage, research-to-schema, and skill synthesis cadence
0.12.0 2026-08-31 August 2026 refresh. Model lineup and all pricing re-verified against primary vendor pages (GPT-5.6 Sol/Terra/Luna, Claude Opus 5 and Sonnet 5, Gemini 3.7 Flash, Grok 4.6, Muse Glimmer). New §2.14 health/money/travel, §3.10 adoption counter-evidence, §4.6.3 discovery of AI prompts, §4.10.1 Claude for Legal, §4.11 litigation workflows. Rule 1.5 billing, privilege-vs-confidentiality, and audit-log-as-ESI added to §4.6, four-part citation check, standing-order step, and Rule 3.3 remediation protocol added to §4.7. Corrected a fabricated Bunce citation carried since Bundle B. Sources 239-283
0.13.0 2026-08-31 Adversarial-review remediation and operator-pattern refresh. Corrected SB 574, Heppner, Doiban, Couvrette, prompt-discovery, FRE 707, Rule 3.3, ESI, security survey, EchoLeak, research-scope, and pricing claims. Added prompt compilation, Ralph backpressure, the functional triad, claim-evidence ledgers, §5.10, and Operator Field Notes. Sources 284-290
0.14.0 2026-08-31 Operator-pattern expansion from a six-lane research fleet with independent verification and local tests. Measured evidence added for disagreement routing and against agreement-as-verification (§5.6.1), completion-contract convergence typology and referee separation (§5.10), stall detection with external progress metrics (§5.4), compaction decay and context placement (§5.5), knowledge-lint boundaries and hash provenance (§5.10), memory-file poisoning hazard (§5.8). METR follow-up added to Part 1 and §3.3. Honest efficacy labels on the prompt compiler and cost per accepted artifact. Companion expanded to 17 pattern entries. Sources 291-332
0.14.1 2026-09-01 Patch. Bibliography hanging indent widened so three-digit reference numbers align with the rest of the list, the HUB International savings figure labeled vendor-published, and the Whiting caveat opener recast to lead with the point

This guide is a living draft. Model names, prices, and product details carry an as-of date because they move every few months. Re-verify anything time-sensitive before relying on it.

Version cadence. This guide follows semantic versioning, and the release triggers are fixed so the credibility of the as-of dates does not depend on guesswork. A patch release (0.0.x) goes out when a listed price moves by more than 10%, when a named product is renamed or retired, or when a citation breaks. A minor release (0.x.0) goes out when a major lab ships a new flagship model that changes the rankings in Part 6, or when a full Part gains or loses a section. A major release (x.0.0) is reserved for a structural rewrite or a change in who the guide is for. Part 6 is the designated maintenance target and may be patched on its own while the rest of the guide holds stable. Routine price and ranking drift that crosses no threshold is batched into the next scheduled review rather than shipped piecemeal.

Review cadence. A full review runs quarterly regardless of whether a trigger has fired, so nothing time-sensitive ages out unnoticed. The review re-verifies every price, model ranking, and product name against its source, checks that each [N] footnote still resolves, and confirms the as-of dates in Part 1, Part 6, and Appendix C. Between quarterly reviews, the triggers above govern: a price move or a renamed product ships as a patch when it is caught, without waiting for the calendar. Anything flagged during a review but not yet resolved is recorded in this changelog so the next pass picks it up.

Appendix F: Source bibliography

The master bibliography is research/_sources.md, a single deduplicated list of 332 numbered entries compiled across the research bundles. Every numbered footnote in this guide, written as [N], keys to that file. The rendered HTML and DOCX editions inline this list as a navigable References section, so each [N] links straight to its source.

Sources are grouped there by category for browsing, then numbered consecutively across the whole list. Entries that churn, such as model versions and pricing, are flagged as time-sensitive and should be re-verified before publication. The bibliography is maintained separately so footnote numbers stay stable as the guide grows.

Social posts and practitioner threads have a narrower role. They can support workflow patterns, vocabulary, and emerging user behavior. They should not be used as the sole source for model performance, pricing, legal obligations, security controls, product availability, or regulatory claims. Those need a primary source or a clearly labeled secondary source.

Appendix G: Context Management Checklist

Use this before a long chat, a new coding-agent run, or a handoff to another model. Section 5.5 explains the reasoning behind each item.

References

  1. Anthropic, "Introducing Claude Opus 4.8," https://www.anthropic.com/news/claude-opus-4-8, May 28, 2026. Cited by: A[1], B[4] (multiple sub-models). [time-sensitive]
  2. Anthropic, "Introducing Claude Opus 4.5," https://www.anthropic.com/news/claude-opus-4-5, Nov 24, 2025; Anthropic system card and platform release notes. Cited by: A[2]. [time-sensitive]
  3. Anthropic, "Claude 3.7 Sonnet and Claude Code," https://www.anthropic.com/news/claude-3-7-sonnet, Feb 24, 2025. Cited by: B[Gemini sub-bundle, sources 2, 3]. [time-sensitive]
  4. Anthropic, "The Claude 3 Model Family: Opus, Sonnet, Haiku," https://www.anthropic.com/news/claude-3-family, Mar 4, 2024. Cited by: B[DeepSeek sub-bundle, source 5].
  5. Anthropic, "Claude 3.5 Sonnet and Artifacts," https://www.anthropic.com/news/claude-3-5-sonnet, Jun 20, 2024. Cited by: B[DeepSeek sub-bundle, source 6].
  6. Anthropic, "Introducing the Model Context Protocol," https://www.anthropic.com/news/model-context-protocol, Nov 25, 2024. Cited by: B[DeepSeek sub-bundle, source 32]. [time-sensitive]
  7. Anthropic, "Many-shot jailbreaking," https://www.anthropic.com/research/many-shot-jailbreaking, May 2024. Cited by: B[DeepSeek sub-bundle, source 61].
  8. OpenAI, "Introducing GPT-5.5," https://openai.com/index/introducing-gpt-5-5, 2026. Cited by: A[3]. [time-sensitive]
  9. OpenAI, "Introducing GPT-5.2" and "Introducing GPT-5.2-Codex," https://openai.com/index/introducing-gpt-5-2 / https://openai.com/index/introducing-gpt-5-2-codex, Dec 11, 2025. Cited by: A[4]. [time-sensitive]
  10. OpenAI, "Introducing ChatGPT Pro," https://openai.com/index/introducing-chatgpt-pro/, Dec 5, 2024. Cited by: B[DeepSeek sub-bundle, source 3; Gemini sub-bundle, source 2].
  11. OpenAI, "Memory and new controls for ChatGPT," https://openai.com/index/memory-and-new-controls-for-chatgpt/, Apr 29, 2024. Cited by: B[DeepSeek sub-bundle, source 4].
  12. OpenAI, "Hello GPT-4o," https://openai.com/index/hello-gpt-4o/, May 13, 2024. Cited by: B[DeepSeek sub-bundle, source 1; Copilot sub-bundle, source 1].
  13. OpenAI, "GPT-4o System Card," https://openai.com/index/gpt-4o-system-card/, May 2024. Cited by: B[DeepSeek sub-bundle, source 2].
  14. OpenAI, "Introducing OpenAI o1-preview," https://openai.com/index/introducing-openai-o1-preview/, Sep 12, 2024. Cited by: B[DeepSeek sub-bundle, source 22].
  15. OpenAI, "GPT-4," https://openai.com/index/gpt-4/, Mar 14, 2023. Cited by: B[DeepSeek sub-bundle, source 66].
  16. OpenAI, "DevDay announcements," https://openai.com/index/new-models-and-developer-products-announced-at-devday/, Nov 6, 2023. Cited by: B[DeepSeek sub-bundle, source 68].
  17. OpenAI Help Center, "About ChatGPT Pro tiers," https://help.openai.com/en/articles/9793128-about-chatgpt-pro-plans, accessed 2026. Cited by: B[TechCrunch sub-bundle, source 3]. [time-sensitive]
  18. OpenAI Help Center, "Best practices for prompt engineering with the OpenAI API," https://help.openai.com/en/articles/6654000-best-practices-for-prompt-engineering-with-openai-api. Cited by: B[Meta sub-bundle, source 12].
  19. OpenAI, "Prompt Engineering Guide," https://platform.openai.com/docs/guides/prompt-engineering, updated 2024. Cited by: B[DeepSeek sub-bundle, source 20].
  20. OpenAI, "ChatGPT Enterprise," https://openai.com/enterprise, 2024. Cited by: B[DeepSeek sub-bundle, source 38].
  21. OpenAI, "ChatGPT Team," https://openai.com/chatgpt/team/, 2024. Cited by: B[DeepSeek sub-bundle, source 48].
  22. OpenAI, "Unlocking Economic Opportunity: ChatGPT-Powered Productivity," https://openai.com [URL not captured in bundle — verify against openai.com], Jul 2025. Cited by: B[ChatGPT sub-bundle, block 11]. [time-sensitive]
  23. OpenAI, "Morgan Stanley case study," https://openai.com/customer-stories/morgan-stanley, 2024. Cited by: B[DeepSeek sub-bundle, source 55].
  24. OpenAI Research Team, "Learning to Reason with LLMs," https://openai.com/index/learning-to-reason-with-llms/, Sep 12, 2024 (updated Jan 2025 with o3 specs). Cited by: B[Gemini sub-bundle, sources 1, 2].
  25. Google, "Gemini 3.1 Pro: Announcing our latest Gemini AI model," https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/, Feb 19, 2026. Cited by: B[TechCrunch sub-bundle, source 5; Meta sub-bundle, source 5]. [time-sensitive]
  26. Google, "Gemini 3" (blog.google/products/gemini/gemini-3), Nov 18, 2025; Google DeepMind Gemini 3 Pro and 3.1 Pro model cards. (A: "Google, 'Gemini 3'"). Also: https://deepmind.google/technologies/gemini/ (updated Mar 2026). Cited by: A[6], B[Gemini sub-bundle, sources 3, 5]. [time-sensitive]
  27. Google, "Google AI Pro & Ultra — get access to Gemini 3.1 Pro & more," https://gemini.google/subscriptions/, accessed 2026. Cited by: B[TechCrunch sub-bundle, source 6]. [time-sensitive]
  28. Google, "Gemini prompt design strategies," https://ai.google.dev/gemini-api/docs/prompting-strategies, 2024. Cited by: B[DeepSeek sub-bundle, source 21; Meta sub-bundle, source 13].
  29. Google, "Google One AI Premium Plan Options," https://one.google.com/explore/plan/all-inclusive, verified active May 2026. Cited by: B[Gemini sub-bundle, source 6]. [time-sensitive]
  30. Google DeepMind, "Gemini 1.5 Pro," https://blog.google/technology/ai/google-gemini-next-generation-model-february-2024/, Feb 15, 2024. Cited by: B[DeepSeek sub-bundle, source 7].
  31. Google, "Everything new in our Google AI subscriptions, fresh from I/O 2026," https://ai.google [URL not fully captured in bundle — verify against ai.google or Google I/O 2026 blog], May 18, 2026. Cited by: B[Grok sub-bundle pricing section]. [time-sensitive]
  32. xAI, "Announcing Grok-1.5," https://x.ai/blog/grok-1.5, Apr 12, 2024. Cited by: B[DeepSeek sub-bundle, source 11].
  33. xAI, "Grok 3 Architecture and Real-Time Supercomputing," https://x.ai/blog/grok-3, Feb 2025. [UNSOURCED — verify per bundle B Gemini sub-bundle note]. Cited by: B[Gemini sub-bundle, sources 1, 7].
  34. Meta AI, "Introducing Llama 3.1," https://ai.meta.com/blog/meta-llama-3-1/, Jul 23, 2024. Cited by: B[DeepSeek sub-bundle, source 9].
  35. Meta AI, "The Llama 3 Herd of Models," https://arxiv.org/abs/2407.21783, arXiv Jul 2024 (training infrastructure detail). Cited by: B[DeepSeek sub-bundle, source 10].
  36. Meta AI, "Meta Llama 3," https://ai.meta.com/blog/meta-llama-3/, Apr 18, 2024. Cited by: B[DeepSeek sub-bundle, source 69].
  37. Meta AI, "Llama 2," https://ai.meta.com/blog/llama-2/, Jul 18, 2023. Cited by: B[DeepSeek sub-bundle, source 67].
  38. Meta AI, "The Llama 4 herd," https://ai.meta.com/blog/llama-4-multimodal-intelligence, Apr 5, 2025. Cited by: A[8]. [breadth]
  39. Meta AI, "Llama 3.3: Expanding the Open Weights Frontier," https://ai.meta.com/blog/meta-llama-3-3/, Dec 6, 2024. Cited by: B[Gemini sub-bundle, source 8].
  40. Meta AI, "Llama 3.2: Open-weights Multimodal Models for Edge Devices," https://arxiv.org/abs/2409.11340, Sep 2024. Cited by: B[Gemini sub-bundle, source 9].
  41. Reuters, "Meta unveils first AI model from costly superintelligence team" (Muse Spark), https://www.reuters.com/sustainability/sustainable-finance-reporting/meta-unveils-first-ai-model-superintelligence-team-2026-04-08/, Apr 8, 2026. Cited by: B[TechCrunch sub-bundle, source 8; Meta sub-bundle, source 8]. [time-sensitive]
  42. Mistral, "Mistral Large," https://mistral.ai/news/mistral-large/, Feb 26, 2024. Cited by: B[DeepSeek sub-bundle, source 13].
  43. Anthropic, "Introducing Claude Opus 4.8" (fast-mode pricing and dynamic workflows detail — same URL as entry 1 but also captured as): https://anthropic.com/claude/opus. Cited by: A[1]. [time-sensitive]
  44. Anthropic Docs, "Pricing," https://www.anthropic.com/pricing, accessed May 31, 2026. Cited by: B[Gemini sub-bundle, source 4; Perplexity sub-bundle]. [time-sensitive]
  45. TechCrunch, Lucas Ropek, "OpenAI releases GPT-5.5, bringing company one step closer to an AI 'super app'," https://techcrunch.com/2026/04/23/openai-chatgpt-gpt-5-5-ai-model-superapp/, Apr 23, 2026. Cited by: B[TechCrunch sub-bundle, source 1; Meta sub-bundle, source 1]. [time-sensitive]
  46. TechCrunch, Ivan Mehta, "OpenAI releases GPT-5.5 Instant, a new default model for ChatGPT," https://techcrunch.com/2026/05/05/openai-releases-gpt-5-5-instant-a-new-default-model-for-chatgpt/, May 5, 2026. Cited by: B[TechCrunch sub-bundle, source 2; Meta sub-bundle, source 2]. [time-sensitive]
  47. TechCrunch, Maxwell Zeff, "Elon Musk's xAI launches Grok 4 alongside a $300 monthly subscription," https://techcrunch.com/2025/07/09/elon-musks-xai-launches-grok-4-alongside-a-300-monthly-subscription/, Jul 9, 2025. Cited by: B[TechCrunch sub-bundle, source 7; Meta sub-bundle, source 7]. [time-sensitive]
  48. CNBC, "From Llamas to Avocados: Meta's AI strategy," https://cnbc.com [URL not fully captured in bundle], Dec 9, 2025. Cited by: A[9]. [breadth]
  49. tech-insider.org, "ChatGPT vs Copilot" (ChatGPT Enterprise $45–75/seat pricing), https://tech-insider.org [URL not captured in bundle], May 2026. Cited by: A[5]. [secondary; time-sensitive]
  50. The Verge, "Google Gemini's image generation still has accuracy issues," https://www.theverge.com/2024/2/22/24079976/google-gemini-ai-image-generation-accuracy-issues, Feb 2024. Cited by: B[DeepSeek sub-bundle, source 8].
  51. The Verge, "BloombergGPT is obsolete," https://www.theverge.com/2024/5/9/24153215/bloomberg-gpt-ai-llm, 2024. Cited by: B[DeepSeek sub-bundle, source 42].
  52. The Verge, "ChatGPT's memory can be poisoned," https://www.theverge.com/2024/5/21/24161881/chatgpt-memory-poisoning, May 2024. Cited by: B[DeepSeek sub-bundle, source 62].
  53. Bloomberg, "Samsung bans ChatGPT after employees leak code," https://www.bloomberg.com/news/articles/2023-05-02/samsung-bans-chatgpt-after-employees-leak-code, May 2023. Cited by: B[DeepSeek sub-bundle, source 63].
  54. Bloomberg / Entrepreneur (Klarna walk-back), "Klarna CEO: AI customer support produced 'lower quality' work; rehiring human agents," [URL not captured in bundle — verify against Bloomberg May 2025], May 2025. Cited by: A[breadth section]. [breadth]
  55. Andrej Karpathy, X (Twitter) thread on LLM knowledge bases, https://twitter.com/karpathy/status/1757888885807849567 (archived), Feb 2024 (early framing); Apr 3–4, 2026 (viral LLM Wiki posts). Cited by: A[11], B[DeepSeek sub-bundle, source 15; TechCrunch sub-bundle, source 9].
  56. Andrej Karpathy, "llm-wiki.md" GitHub gist, https://gist.github.com/karpathy/llm-wiki.md (exact URL not captured in all bundles; see also community synthesis gist), Apr 4, 2026. Cited by: A[12]. [time-sensitive]
  57. GitHub Gist by deanjstone, "Karpathy's LLM Wiki — A Synthesis of notes, sources, and analysis," https://gist.github.com/deanjstone/98141cb836bb97c555ae9d6ce2484b5f, last updated Apr 19, 2026. Cited by: B[TechCrunch sub-bundle, source 9; Meta sub-bundle, source 9]. [secondary, citing Karpathy's work]
  58. Simon Willison, "Codex CLI 0.128.0 adds /goal," https://simonwillison.net/2026/Apr/30/codex-goals, Apr 30, 2026. Cited by: A[16]. [primary-adjacent]
  59. Simon Willison's Weblog, various posts on AI wrapping, security, and prompt injection, https://simonwillison.net/. Cited by: A[breadth], B[DeepSeek sub-bundle, source 43].
  60. GitHub, openai/codex, "Release 0.128.0," https://github.com/openai/codex/releases/tag/rust-v0.128.0. Cited by: B[TechCrunch sub-bundle, source 10; Meta sub-bundle, source 10]. [time-sensitive]
  61. ralphable.com, "Codex /goal Command — Built-in Ralph Loop," https://ralphable.com [URL not captured in bundle], Apr 30, 2026; MindStudio, "Codex /goal — 14-hour autonomous task," [URL not captured in bundle]. Cited by: A[14]. [secondary]
  62. Thomas Wiegold, "The Ralph Loop: How Recursive AI Agents Actually Work," https://thomas-wiegold.com [URL not captured in bundle], May 2026 (quotes Geoffrey Huntley). Cited by: A[15]. [secondary]
  63. Bloomberg Law, Eric Killelea and Roy Strom, "Kirkland & Ellis Investing $500 Million to Build AI Platform," https://news.bloomberglaw.com/business-and-practice/kirkland-ellis-investing-500-million-to-build-ai-platform, May 28, 2026. (FT reported first; BLaw same day.) Cited by: A[29], B[Copilot sub-bundle, source 12; Meta sub-bundle, source 14]. [time-sensitive]
  64. Reuters, "Kirkland & Ellis to invest $500 million in AI platform," https://www.reuters.com [URL not fully captured in bundle — verify], May 28, 2026. Cited by: B[ChatGPT sub-bundle, block 8]. [time-sensitive]
  65. The Global Legal Post, "Kirkland to invest $500m building own AI platform," https://globallegalpost.com [URL not captured in bundle], May 2026; Convergences (Substack), "The $500 Million Answer." Cited by: A[30]. [secondary]
  66. Lawfuel, "BigLaw's AI Arms Race — Kirkland & Ellis $500 Million," https://lawfuel.com [URL not captured in bundle], May 2026 (Jon Ballis "floor" quote). Cited by: A[31]. [secondary]
  67. Harvey.ai blog, "Harvey Raises at $11 Billion Valuation to Scale Agents Across Law Firms and Enterprises," https://www.harvey.ai/blog/harvey-raises-at-dollar11-billion-valuation-to-scale-agents-across-law-firms-and-enterprises, Mar 25–26, 2026. Cited by: A[34], B[TechCrunch sub-bundle, source 13 (Harvey blog); ChatGPT sub-bundle, block 8]. [time-sensitive]
  68. Reuters, "Legal software firm Harvey valued at $11 billion in latest funding round," https://www.reuters.com/technology/legal-software-firm-harvey-valued-11-billion-latest-funding-round-2026-03-25/, Mar 25, 2026. Cited by: B[Meta sub-bundle, source 15]. [time-sensitive]
  69. Sacra, "Harvey revenue, valuation & funding," https://sacra.com/c/harvey/, updated ~Apr 27, 2026 (BigLaw Bench "scrapped model" + commoditization characterization). Cited by: A[35], B[BigLaw sub-bundle, source 1; Build/Buy sub-bundle, source 1]. [analyst/secondary; time-sensitive]
  70. CNBC and Bloomberg, Harvey $200M fundraise coverage, https://www.cnbc.com [URL not captured in bundle], Mar 25, 2026. Cited by: A[34], B[Copilot sub-bundle, source 13].
  71. Harvey AI, TechCrunch, "Harvey raises $100M at $1.5B valuation," https://techcrunch.com/2024/07/31/harvey-ai-raises-100m/, Jul 2024. (Earlier fundraise, context for scale.) Cited by: B[DeepSeek sub-bundle, source 33].
  72. Sacra Research, "Hebbia revenue, valuation & funding," https://sacra.com/c/hebbia/, May 25, 2026 (Hebbia $700M Series B, FlashDocs acquisition). Cited by: B[BigLaw sub-bundle, source 4]. [time-sensitive]
  73. Thomson Reuters, "Legal AI in 2026: Why CoCounsel thrives while others fold," https://legal.thomsonreuters.com/blog/legal-ai-in-2026-why-cocounsel-thrives-while-others-fold/, Mar 19, 2026. Cited by: B[BigLaw sub-bundle, source 5]. [time-sensitive]
  74. Thomson Reuters, "CoCounsel Legal Reimagined" (Apr 2026 agentic rebuild announcement), https://www.thomsonreuters.com [URL not fully captured in bundle — verify], Apr 22, 2026. Cited by: B[ChatGPT sub-bundle, block 8]. [time-sensitive]
  75. Thomson Reuters, acquisition of Casetext press release, https://www.thomsonreuters.com/en/press-releases/2023/june/thomson-reuters-acquires-casetext.html, Jun 27, 2023. Cited by: B[DeepSeek sub-bundle, source 37].
  76. Freshfields press release, "Freshfields–Anthropic partnership (Claude to 5,700 employees)," https://freshfields.com [URL not captured in bundle], Apr 23, 2026. Cited by: A[breadth section]. [breadth; time-sensitive]
  77. Spellbook, "Which Law Firms Use AI? Case Studies From BigLaw to Solo Practices," https://spellbook.com/learn/law-firms-using-ai, May 11, 2026. Cited by: B[BigLaw sub-bundle, source 6]. [secondary; time-sensitive]
  78. Davis Polk, "Artificial Intelligence" practice page, https://www.davispolk.com/artificial-intelligence, and the "How generative AI is impacting the legal industry" webinar (Dec 11, 2024). Named AI-practice voices include partners Matthew Bacal (corporate lead), James "Jamie" Haldin, and David Lisson, plus Director of Knowledge Management Rebecca Yeargan. No flagship internal AI build has been publicly announced as of mid-2026, unlike peer firms that publicized Harvey or CoCounsel rollouts. Source upgraded 2026-06-08 (real URL + named partners) during the §4.9 firm-accuracy pass. Cited by: A[32], guide.md §4.9. [breadth]
  79. LawSites / Bob Ambrogi, "LexisNexis announces Protégé General AI," https://lawnext.com [URL not captured in bundle; also artificiallawyer.com], Dec 10, 2025. Cited by: B[ChatGPT sub-bundle, block 8]. [breadth]
  80. Above the Law, "Is the legal AI bubble about to burst?" https://abovethelaw.com [URL not captured in bundle], Mar 16, 2026. Cited by: B[ChatGPT sub-bundle, block 8]. [secondary]
  81. Reddit BigLaw Forum, "Harvey valuation at $11 billion," https://www.reddit.com/r/biglaw/comments/1s3hfzq/harvey_valuation_at_11_billion/, Mar 26, 2026. Cited by: B[BigLaw sub-bundle, source 2]. [secondary]
  82. Latham & Watkins, press release, "LathamWatches," https://www.lw.com/news/lathamwatches, Aug 2023. Cited by: B[DeepSeek sub-bundle, source 34].
  83. American Bar Association, "Formal Opinion 512: Generative Artificial Intelligence Tools," https://www.americanbar.org/groups/professional_responsibility/publications/ethics_opinions/aba-formal-opinion-512/, Jul 29, 2024. (Bundle B DeepSeek cites "Oct 2024" — exact date per ABA is Jul 29, 2024.) Cited by: A[36], B[DeepSeek sub-bundle, source 44; Copilot sub-bundle, source 15; ChatGPT sub-bundle, block 10].
  84. Frankfurt Kurnit (technologylaw.fkks.com), Kristen Niven, "ABA Issues Comprehensive Formal Ethics Opinion on Lawyers' Use of Generative AI," https://technologylaw.fkks.com [URL not captured in bundle], 2024; americanbar.org Business Law Today (Oct 2024). Cited by: A[37]. [secondary]
  85. State Bar of California, "Practical Guidance for the Use of Generative Artificial Intelligence," https://www.calbar.ca.gov/ethics, Nov 2023 (updated 2024/2026). Cited by: A[38], B[DeepSeek sub-bundle, source 45; Perplexity sub-bundle, block 13]. [time-sensitive]
  86. New York City Bar Association, "Current Ethics Opinions and Reports Related to Generative AI" (PDF), https://nycbar.org [URL not captured in bundle], May 2025. Cited by: A[38].
  87. New York State Bar Association, "AI Task Force Report," https://nysba.org/ai-task-force-report/, 2024. Cited by: B[DeepSeek sub-bundle, source 46].
  88. Mata v. Avianca, S.D.N.Y., sanctions order, https://www.courtlistener.com/docket/63107780/mata-v-avianca-inc/, Jun 2023. Cited by: A[39], B[DeepSeek sub-bundle, source 47; ChatGPT sub-bundle, block 10].
  89. Damien Charlotin, AI Hallucination Cases Database, https://damiencharlotin.com/hallucinations (1,227+ cases globally by early 2026). Cited by: A[40]. [time-sensitive]
  90. PlatinumIDS Blog, "1,227 Fabricated Citations and Counting," https://platinumids.com [URL not captured in bundle], 2026; Bloomberg Law, "AI-Faked Cases Become Core Issue Irritating Overworked Judges." Cited by: A[40]. [secondary; time-sensitive]
  91. Sterne Kessler, "AI Hallucinations in Court Filings — 2025 Review," https://sternekessler.com [URL not captured in bundle], 2025/2026; getvoibe.com "AI Hallucinations in Law Firms (2026)." Cited by: A[41]. [secondary]
  92. thelegalprompts.com, "AI Hallucinations in Legal Work" (Mata v. Avianca summary; Crabill 90-day suspension), https://thelegalprompts.com [URL not captured in bundle], 2026. Cited by: A[39, 42]. [secondary]
  93. GC AI, "AI Legal Ethics in 2026: 6 Cases, 4 Rules, 1 Policy Template," https://gc.ai [URL not captured in bundle], Jan 2026. Cited by: B[ChatGPT sub-bundle, block 10]. [secondary]
  94. JD Supra, "Mata v. Avianca and the Risk of AI Hallucinations," https://jdsupra.com [URL not captured in bundle], Aug 2023 (updated 2024). Cited by: B[ChatGPT sub-bundle, block 10].
  95. California Legislature, Senate Bill 574 (Umberg), "Attorney and Arbitrator Use of Generative Artificial Intelligence," official text and status at https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260SB574 and https://leginfo.legislature.ca.gov/faces/billStatusClient.xhtml?bill_id=202520260SB574. The original research bundle incorrectly described the bill as passed in January 2026. As of August 31, 2026, it remained active in the Assembly floor process and was not law. Source [252] contains the current operative-text summary. Cited by: B[BigLaw-legal sub-bundle, block 10]. [primary; legislature; time-sensitive; corrected]
  96. Bunce v. Visual Technology Innovations, No. 2:23-cv-01740 (E.D. Pa., Judge Kai N. Scott), Rule 11 AI-sanctions memorandum and order entered April 20, 2026 at docket entries 225–226, https://www.courtlistener.com/docket/67332015/. The original research bundle supplied a false reporter number and date. Source [278] preserves the correction record and sanctions details. Cited by: B[BigLaw-legal sub-bundle, block 10]. [primary; docket; corrected]
  97. Florida Bar, "Opinion 24-1," Jan 2024 (AI use duties). Cited by: B[ChatGPT sub-bundle, block 10]. [URL not captured in bundle — verify via floridabar.org].
  98. METR (Becker, Rush, Barnes, Rein), "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," https://metr.org/blog/2025-07-10-measuring-the-impact-of-early-2025-ai-on-experienced-open-source-developer-productivity and arXiv:2507.09089, Jul 10, 2025. Randomized trial: 16 experienced open-source developers, 246 tasks in familiar repositories, early-2025 tools measured 19% slower (CI +2% to +39%) while developers forecast a 24% speedup beforehand and self-assessed 20% afterward. Cited by: A[21]; guide.md Part 1 and §3.3.
  99. METR, "We are Changing our Developer Productivity Experiment Design," https://metr.org/blog/2026-02-24-uplift-update, Feb 24, 2026, re-verified August 31, 2026. Late-2025 data: the returning original cohort measured a speedup of -18% in change-of-completion-time terms, that is roughly 18% faster (CI -38% to +9%), and newly recruited developers about 4% faster (CI -15% to +9%), both intervals crossing zero. METR calls this new data "an unreliable signal," primarily because developers increasingly declined to participate rather than work without AI, a bias METR says pushes its speedup estimate down, making the estimate a likely lower bound. Its chart caption: "Late-2025 AI likely accelerated open-source developers, but selection effects obscure the true speedup." Also reports 30% to 50% of developers withholding some tasks they preferred doing with AI. Cited by: A[22]; guide.md Part 1 and §3.3. [first-party follow-up; self-described unreliable signal]
  100. Li, Xiaofan, and Xing Gao, "A First Look at the Security Issues in the Model Context Protocol Ecosystem," arXiv:2510.16558 (accepted DSN 2026), v1 Oct 18, 2025, v2 Apr 27, 2026. 67,057 public MCP servers counted across six registries (data collected late June–early July 2025). Cited by: A[25], guide.md §5.8 / Appendix. Note: attribution corrected 2026-06-08 from an earlier "Hu et al." label — the paper has no author named Hu; lead authors are Li and Gao.
  101. Wu, Wu, and Zou (Stanford), "FineTuneBench," arXiv:2411.05059 (~37% generalization for new knowledge via fine-tuning). Cited by: A[breadth section]. [breadth]
  102. Liu et al., "Lost in the Middle: How Language Models Use Long Contexts," arXiv:2307.03172, Nov 2023. Cited by: B[DeepSeek sub-bundle, source 23].
  103. Cognition, "Introducing Devin, the first AI software engineer," https://www.cognition.ai/blog/introducing-devin, Mar 12, 2024. Cited by: B[DeepSeek sub-bundle, source 16].
  104. Dell'Acqua, F., et al., "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality," Harvard Business School Working Paper 24-013, https://www.hbs.edu/faculty/Pages/item.aspx?num=64700 (SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321), Sep 2023. The BCG-consultant field experiment: on tasks inside the model's frontier, consultants using GPT-4 completed 12.2% more tasks, 25.1% faster, with results rated >40% higher in quality; on tasks outside it they performed worse. Cited by: guide.md §3.3. [academic; primary]
  105. McKinsey, "The economic potential of generative AI," https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/the-economic-potential-of-generative-ai, Jul 2024. Cited by: B[DeepSeek sub-bundle, source 52].
  106. MIT NANDA (lead author Aditya Challapally), "The GenAI Divide: State of AI in Business 2025," https://mlq.ai [PDF mirror; verify against nanda.mit.edu], Jul 2025. (95% of enterprise gen-AI pilots show no measurable P&L impact; methodology noted as preliminary.) Cited by: A[breadth section], B[Build/Buy sub-bundle, source 8 indirectly]. [breadth]
  107. Bloomberg, "BloombergGPT: A Large Language Model for Finance," https://www.bloomberg.com/company/press/bloomberggpt-50-billion-parameter-llm/, Mar 2023. Cited by: B[DeepSeek sub-bundle, source 40].
  108. OpenAI, "GPT-4 Turbo long context evaluation," https://openai.com/index/new-models-and-developer-products-announced-at-devday/, Nov 2023. Cited by: B[DeepSeek sub-bundle, source 24].
  109. Sequoia Capital, "The Rise of the Agentic Loop in Enterprise Workflows," https://www.sequoiacapital.com/article/agentic-loop-enterprise/, Sep 18, 2025. [URL stability noted as UNSOURCED in bundle — verify]. Cited by: B[Gemini sub-bundle, source 11]. [secondary]
  110. Apple, "Apple Intelligence," https://www.apple.com/newsroom/2024/06/introducing-apple-intelligence-for-iphone-ipad-and-mac/, announced Jun 2024. Cited by: B[DeepSeek sub-bundle, source 18].
  111. National Security Agency, "Model Context Protocol (MCP)" Cybersecurity Information Sheet, https://nsa.gov [URL not captured in bundle — verify via nsa.gov], May 2026. Cited by: A[28]. [time-sensitive]
  112. NIST, "AI Risk Management Framework 1.0," https://www.nist.gov/itl/ai-risk-management-framework, Jan 26, 2023. Cited by: A[breadth section], B[DeepSeek sub-bundle, source 65; Copilot sub-bundle, source 15].
  113. NIST, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile" (NIST AI 600-1), https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf, Jul 26, 2024. Cited by: A[breadth section]. [breadth]
  114. OWASP, "OWASP Top 10 for LLM Applications" (2025 list, release v4.2.0a), https://owasp.org/www-project-top-10-for-large-language-model-applications/ (also https://genai.owasp.org/llm-top-10), 2025. Cited by: A[breadth section], B[DeepSeek sub-bundle, source 60]. [time-sensitive]
  115. Anthropic, "Effective context engineering for AI agents," https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents (also cited as /blog/context-engineering in some bundles), Sep 29, 2025. Cited by: A[17], B[TechCrunch sub-bundle, source 11 (with URL); Meta sub-bundle, source 11; Copilot sub-bundle; Perplexity sub-bundle; Gemini sub-bundle, source 7; Grok sub-bundle].
  116. Anthropic, "Effective harnesses for long-running agents," https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents, 2025/2026. Cited by: A[19].
  117. Anthropic, "Scaling Managed Agents: Decoupling the Brain from the Hands," https://www.anthropic.com/engineering/managed-agents, Apr 8, 2026. Cited by: B[Gemini sub-bundle, source 10]. [time-sensitive]
  118. Isabella He, "Context engineering: memory, compaction, and tool clearing," Claude Cookbook, https://platform.claude.com/cookbook/tool-use-context-engineering-context-engineering-tools, Mar 20, 2026. Cited by: B[Gemini sub-bundle, source 8]. [time-sensitive]
  119. Anthropic, "How to give Claude a persona with a system prompt," https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/system-prompts, Sep 2024. Cited by: B[DeepSeek sub-bundle, source 19].
  120. Anthropic, "Mitigating prompt injection," https://docs.anthropic.com/en/docs/security/prompt-injection, 2024. Cited by: B[DeepSeek sub-bundle, source 64].
  121. Anthropic Docs, "MCP connector," https://docs.anthropic.com [URL — verify current beta header mcp-client-2025-11-20], 2025–2026. Cited by: B[Perplexity sub-bundle, block 7]. [time-sensitive]
  122. Google Cloud, "What Is AI context engineering?" https://cloud.google.com [URL not fully captured in bundle], last updated Apr 23, 2026. Cited by: B[Grok sub-bundle, Perplexity sub-bundle, block 4]. [time-sensitive]
  123. Atlan, "Harness Engineering vs Prompt Engineering vs Context Engineering," https://atlan.com [URL not captured in bundle], Mar 2026. Cited by: B[ChatGPT sub-bundle, block 4]. [secondary]
  124. AI Agents Plus, "AI Context Window Management Techniques," https://aiagentsplus.com [URL not captured in bundle], Mar 8, 2026. Cited by: B[ChatGPT sub-bundle, blocks 4, 5]. [secondary]
  125. Andrej Karpathy, "Intro to Large Language Models: LLM OS," https://karpathy.ai/, Nov 2023 (expanded to LLM Wiki framework through mid-2025 via public presentations). Cited by: B[Gemini sub-bundle, source 10]. [direct URL stability noted as UNSOURCED in bundle — verify]
  126. MarkTechPost, "A Guide for Effective Context Engineering for AI Agents," https://marktechpost.com [URL not captured in bundle], Oct 20, 2025 (summarizing Anthropic; "context rot"). Cited by: A[18]. [secondary]
  127. Data Science Dojo, "LLM Wiki by Andrej Karpathy: Tutorial," https://datasciencedojo.com/blog/llm-wiki-tutorial, Apr 2026. Cited by: A[12]. [secondary]
  128. Starmorph, "How to Build Karpathy's LLM Wiki," https://blog.starmorph.com [URL not captured in bundle], Apr 2026 (quotes Karpathy). Cited by: A[13]. [secondary]
  129. Urvil Joshi, "Andrej Karpathy's LLM Wiki," https://medium.com [URL not captured in bundle], Apr 2026; AI Builder Club, "Karpathy's LLM Wiki," https://aibuilderclub.com [URL not captured in bundle], Apr 2026. Cited by: A[11]. [secondary, citing Karpathy's X post + gist]
  130. OpenAI, "Understanding prompt injections: a frontier security challenge," https://openai.com [URL not captured in bundle — verify via openai.com/blog], Nov 6, 2025. Cited by: B[Perplexity sub-bundle, block 13].
  131. Securance, "Prompt injection: the OWASP #1 AI threat in 2026," https://securance.com [URL not captured in bundle], Apr 14, 2026. Cited by: B[Perplexity sub-bundle, security section]. [secondary]
  132. Airia, "AI Security in 2026: Prompt Injection, the Lethal Trifecta, and How to Defend," https://airia.com [URL not captured in bundle], May 26, 2026. Cited by: B[Perplexity sub-bundle, security section]. [secondary; time-sensitive]
  133. GuruSup, "AI Models in 2026," https://gurusup.com/blog/ai-comparisons, May 2, 2026; Pluralsight, "Best AI models in 2026." Cited by: A[10]. [secondary; time-sensitive]
  134. Red Hat, "Model Context Protocol (MCP): Understanding security risks and controls," https://redhat.com [URL not captured in bundle], 2025/2026. Cited by: A[23].
  135. authzed.com, "A Timeline of Model Context Protocol (MCP) Security Breaches," https://authzed.com [URL not captured in bundle], updated Apr 2026. Cited by: A[24]. [secondary]
  136. CIO, "Why Model Context Protocol is suddenly on every executive agenda," https://cio.com [URL not captured in bundle], 2026. Cited by: A[26]. [secondary]
  137. The Hacker News, "Anthropic MCP Design Vulnerability Enables RCE across 7,000+ servers / 150M+ downloads," https://thehackernews.com/2026/04 [exact URL not captured in bundle], Apr 2026 (OX Security research credited). Cited by: A[27]. [secondary, primary research credited]
  138. lushbinary.com, "AI Coding Agents 2026 — Pricing & Features," https://lushbinary.com [URL not captured in bundle], updated May 20, 2026; builder.io "Claude Code vs Cursor" 2026; dev.to "Cursor vs Windsurf vs Claude Code." Cited by: A[20]. [secondary; time-sensitive]
  139. CloudZero FinOps Lab, "Claude Code Pricing In 2026: Plans, Token Costs, And What It Actually Costs to Use," https://www.cloudzero.com/blog/claude-code-pricing/, Feb 28, 2026. Cited by: B[Gemini sub-bundle coding section, source 1]. [secondary; time-sensitive]
  140. TLDL Developer Resources, "AI Coding Tools Compared (2026): Cursor vs Claude Code vs Copilot — Benchmarks & Pricing," https://www.tldl.io/resources/ai-coding-tools-2026, Mar 15, 2026. Cited by: B[Gemini sub-bundle coding section, source 3; ChatGPT sub-bundle, block 6]. [secondary; time-sensitive]
  141. Justin McKelvey, "Cursor vs Windsurf 2026: Honest Comparison (I Used Both)," https://justinmckelvey.com/blog/cursor-vs-windsurf, May 24, 2026. Cited by: B[Gemini sub-bundle coding section, source 2]. [secondary; time-sensitive]
  142. Digital Applied Research, "MCP Adoption Statistics 2026: Model Context Protocol," https://www.digitalapplied.com/blog/mcp-adoption-statistics-2026-model-context-protocol, May 26, 2026. Cited by: B[Gemini sub-bundle MCP section, source 7]. [secondary; time-sensitive]
  143. Build Fast With AI, "What Is MCP (Model Context Protocol)? Complete 2026 Guide," https://www.buildfastwithai.com/blogs/what-is-model-context-protocol-mcp, Mar 9, 2026. Cited by: B[Gemini sub-bundle MCP section, source 8]. [secondary]
  144. Scott Blandford / Truthifi, "MCP connectors for Claude, ChatGPT, Perplexity, Grok, Mistral & OpenClaw (2026)," https://truthifi.com/education/mcp-connection-guide, Apr 18, 2026. Cited by: B[Gemini sub-bundle MCP section, source 6]. [secondary; time-sensitive]
  145. Atlan, "What is MCP? Why it matters and how to think about it," https://atlan.com [URL not captured in bundle], May 2026. Cited by: B[ChatGPT sub-bundle, block 7]. [secondary]
  146. Digital Applied, "Enterprise AI Agent Build vs Buy: 2026 Decision," https://www.digitalapplied.com/blog/enterprise-ai-agent-build-vs-buy-2026, May 17, 2026. Cited by: B[Build/Buy sub-bundle, source 7]. [secondary; time-sensitive]
  147. TechAhead, "Build vs Buy vs Partner AI: The Enterprise AI Decision Framework for 2026," https://www.techaheadcorp.com/blog/enterprise-ai-build-vs-buy-vs-partner/, May 4, 2026. Cited by: B[Build/Buy sub-bundle, source 8]. [secondary; time-sensitive]
  148. 8allocate, "Why Build vs Buy AI Is the Wrong Question for Your Product," https://8allocate.com/blog/why-build-vs-buy-ai-is-the-wrong-question-for-your-product/, Apr 7, 2026. Cited by: B[Build/Buy sub-bundle, source 9]. [secondary]
  149. Writer, "Build vs. buy: Scaling agentic AI on a unified platform," https://writer.com/blog/build-vs-buy-generative-ai/, Apr 7, 2026. Cited by: B[Build/Buy sub-bundle, source 10]. [secondary]
  150. Just Think AI, "Google VP Warning: The 2 Types of AI Startups Heading for Failure," https://www.justthink.ai/blog/google-vp-warning-the-2-types-of-ai-startups-heading-for-failure, May 6, 2026. Cited by: B[Build/Buy sub-bundle, source 11]. [secondary]
  151. Easy.bi, "The AI Build vs. Buy Decision: Custom Models vs. API Wrappers," https://easy.bi [URL not captured in bundle], Feb–Mar 2026. Cited by: B[ChatGPT sub-bundle, block 9]. [secondary]
  152. sitepoint.com, "Best Local LLM Models 2026" and "Guide to Local LLMs in 2026," https://sitepoint.com [URL not captured in bundle], 2026; promptquorum.com "Best Local LLMs May 2026." Cited by: A[44]. [secondary; time-sensitive]
  153. codersera.com, "Ollama vs LM Studio vs vLLM vs llama.cpp vs MLX 2026," https://codersera.com [URL not captured in bundle], May 26, 2026; "vLLM vs Ollama vs LM Studio production benchmark" (citing Red Hat Aug 2025); julsimon (Medium), "What to Buy for Local LLMs (April 2026)." Cited by: A[45]. [secondary; time-sensitive]
  154. TechCrunch, "Cursor is rolling out a new kind of agentic coding tool (Automations)," https://techcrunch.com [URL not captured in bundle], Mar 5, 2026. Cited by: B[ChatGPT sub-bundle, block 6]. [time-sensitive]
  155. NxCode, "Windsurf vs Cursor 2026: Which AI IDE Should You Choose?" https://nxcode.io [URL not captured in bundle], Mar 2026. Cited by: B[ChatGPT sub-bundle, block 6]. [secondary; time-sensitive]
  156. VibecodingHub, "Aider Review," https://vibecodinghub.com [URL not captured in bundle], Apr 2026. Cited by: B[ChatGPT sub-bundle, block 6]. [secondary]
  157. Wikipedia, "Codex (AI agent): Architecture and Platform Expansions," https://en.wikipedia.org/wiki/Codex_(AI_agent), last modified Apr 29, 2026. Cited by: B[Gemini sub-bundle coding section, source 5]. [secondary]
  158. OpenAI, "OpenAI named a Leader in enterprise coding agents by Gartner," https://openai.com/index/gartner-2026-agentic-coding-leader/, May 22, 2026. Cited by: B[Gemini sub-bundle coding section, source 4]. [time-sensitive]
  159. McKinsey, "The State of AI in 2025 / AI Productivity Gains and the Performance Paradox," https://www.mckinsey.com/quantumblack [URL not fully captured in bundle], 2025/May 2026. Cited by: A[breadth section], B[ROI sub-bundle, block 11]. [breadth; time-sensitive]
  160. Anthropic / HUB International, "Claude to 20,000+ employees," https://hubinternational.com [URL not captured in bundle], Feb 25, 2026. Cited by: A[breadth section]. [breadth; time-sensitive]
  161. Anthropic, "Customer Stories," https://www.anthropic.com/customers, 2024. Cited by: B[DeepSeek sub-bundle, source 54].
  162. OpenAI, "ChatGPT Pricing," https://openai.com/chatgpt/pricing [URL not captured in bundle — verify], published Feb 24, 2026. Cited by: B[Perplexity sub-bundle, block 11]. [time-sensitive]
  163. Google Support, "Connect your Google apps and third-party data — Gemini Enterprise," https://support.google.com [URL not fully captured — verify via support.google.com], last updated May 22, 2026. Cited by: B[Perplexity sub-bundle, block 7]. [time-sensitive]
  164. Microsoft, "Microsoft 365 Copilot Plans and Pricing," https://www.microsoft.com/en-us/microsoft-365-copilot/pricing, accessed August 31, 2026. Microsoft 365 Copilot Business is $18 per user per month paid yearly, discounted from $21. The offer runs July 1 through December 31, 2026, requires an annual commitment, and applies to the first year. Cited by: guide.md §1.3, §3.5, Appendix C. [primary; vendor; time-sensitive]
  165. GitHub, "Introducing GitHub Copilot Workspace," https://github.blog/2024-04-29-github-copilot-workspace/, Apr 2024. Cited by: B[DeepSeek sub-bundle, source 31].
  166. GitHub, "Copilot pricing," https://github.com/features/copilot/plans (also https://github.com/features/copilot), 2024/2026. Cited by: B[DeepSeek sub-bundle, source 51; Copilot sub-bundle]. [time-sensitive]
  167. Anthropic, "Claude Team plan," https://www.anthropic.com/team, 2024. Cited by: B[DeepSeek sub-bundle, source 49].
  168. Google, "Gemini for Google Workspace," https://workspace.google.com/solutions/ai/, 2024. Cited by: B[DeepSeek sub-bundle, source 50].
  169. LMSYS Chatbot Arena, "Yi-Large ranked #1," https://lmsys.org/blog/2024-05-20-yi-large/, May 2024. Cited by: B[DeepSeek sub-bundle, source 14].
  170. Cursor, "The AI-first Code Editor," https://cursor.sh (also cursor.com), 2024/2026. Cited by: B[DeepSeek sub-bundle, source 26]. [time-sensitive]
  171. Cognition, "Plans and Pricing," https://windsurf.com/pricing, accessed August 31, 2026. The Windsurf URL redirects to Cognition's Devin pricing. Current US figures shown are Pro $20 per month and Teams at $80 per month for the team plan plus $40 per month for each full developer seat. The product page calls the editor Devin Desktop and the documentation identifies Windsurf as a Cognition product. Cited by: guide.md §1.3, §6.2, Appendix C. [primary; vendor; time-sensitive]
  172. Aider, "AI pair programming in your terminal," https://github.com/paul-gauthier/aider, 2024. Cited by: B[DeepSeek sub-bundle, source 28].
  173. Cline, "Autonomous coding agent for VSCode," https://github.com/cline/cline, 2024. Cited by: B[DeepSeek sub-bundle, source 29].
  174. Replit, "Replit Agent," https://replit.com/site/agent, 2024. Cited by: B[DeepSeek sub-bundle, source 30].
  175. Ollama, GitHub repository, https://github.com/ollama/ollama, 2024/2026. Cited by: B[DeepSeek sub-bundle, source 56].
  176. llama.cpp, GitHub repository, https://github.com/ggerganov/llama.cpp, 2024/2026. Cited by: B[DeepSeek sub-bundle, source 57].
  177. LM Studio, https://lmstudio.ai, 2024/2026. Cited by: B[DeepSeek sub-bundle, source 58]. [time-sensitive]
  178. vLLM, GitHub repository, https://github.com/vllm-project/vllm, 2024/2026. Cited by: B[DeepSeek sub-bundle, source 59].
  179. Spellbook, "AI Contract Drafting and Review," https://www.spellbook.legal, 2024. Cited by: B[DeepSeek sub-bundle, source 36].
  180. Notion, "Notion AI," https://www.notion.so/product/ai, 2024. Cited by: B[DeepSeek sub-bundle, source 41].
  181. Meta, "Llama 3.1 fine-tuning guide," https://ai.meta.com/llama/, 2024. Cited by: B[DeepSeek sub-bundle, source 39].
  182. Model Context Protocol specification and GitHub repository, https://modelcontextprotocol.io/specification (and https://modelcontextprotocol.io/), 2025–2026. Cited by: B[DeepSeek sub-bundle noted MCP spec; Grok sub-bundle; Copilot sub-bundle, source 9]. [time-sensitive]
  183. Anthropic, "Building effective agents," https://resources.anthropic.com/building-effective-ai-agents (cited as Claude docs / engineering post), 2025. Cited by: A[breadth footnotes].
  184. Anthropic, "Claude Code docs," https://docs.anthropic.com/en/docs/agents-and-tools/claude-code, 2025/2026. Cited by: B[Gemini sub-bundle primary sources section]. [time-sensitive]
  185. OpenAI, "Business Pricing," https://openai.com/business/pricing/, accessed August 31, 2026. Standard ChatGPT Business seats are $20 per month billed annually or $25 billed monthly. Premium seats are $100 annually or $125 monthly. OpenAI states that business data is not used for training by default. Cited by: guide.md §1.3, §3.5, Appendix C. [primary; vendor; time-sensitive]
  186. Atlan, "What is MCP? Why it matters," plus "Fenxi, 'MCP and why connectors matter,' Jan 2026," and "Digital Applied, 'MCP Adoption Metrics 2026, Mar 2026.'" Cited by: B[ChatGPT sub-bundle, block 7]. (Fenxi URL: https://fenxi.ai [URL not captured]; Digital Applied URL: https://digitalapplied.com/blog/mcp-adoption-statistics-2026.) [secondary]
  187. Anthropic, official ralph-wiggum plugin for Claude Code, https://github.com/anthropics/claude-code/tree/main/plugins/ralph-wiggum (formalizes Geoffrey Huntley's Ralph-loop technique via the Claude Code stop-hook; released Dec 2025). Corroborating coverage: The Register, Jan 27, 2026. Added during the June 2, 2026 Part 1 source-resolution pass. [primary]
  188. Kadavath, S., et al. (Anthropic), "Language Models (Mostly) Know What They Know," arXiv:2207.05221, Jul 2022 (v4 Nov 2022). https://arxiv.org/abs/2207.05221 Cited by: guide.md §2.10 (model self-calibration; confidence-level instruction). [primary]
  189. Brown, T. B., et al. (OpenAI), "Language Models are Few-Shot Learners," NeurIPS 2020, vol. 33, pp. 1877–1901. arXiv:2005.14165. https://arxiv.org/abs/2005.14165 Cited by: guide.md §2.11 (few-shot prompting). [primary]
  190. Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P., "Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity," ACL 2022 (long papers). arXiv:2104.08786. https://arxiv.org/abs/2104.08786 / https://aclanthology.org/2022.acl-long.556/ Cited by: guide.md §2.11 (example ordering; recency-of-example effect). [primary]
  191. Madaan, A., Tandon, N., Gupta, P., et al., "Self-Refine: Iterative Refinement with Self-Feedback," NeurIPS 2023, vol. 36, pp. 46534–46594. arXiv:2303.17651. https://arxiv.org/abs/2303.17651 Cited by: guide.md §2.11 (generate-then-refine loop; ~20% average improvement). [primary]
  192. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., and Zhou, D., "Large Language Models Cannot Self-Correct Reasoning Yet," ICLR 2024. arXiv:2310.01798. https://arxiv.org/abs/2310.01798 Cited by: guide.md §§2.11 and 5.6.1 (limits of self-correction on reasoning tasks); companions/operator-field-notes.md. [primary]
  193. Zheng, M., Pei, J., Logeswaran, L., Lee, M., and Jurgens, D., "When 'A Helpful Assistant' Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models," Findings of EMNLP 2024, pp. 15126–15154. arXiv:2311.10054. https://arxiv.org/abs/2311.10054 Cited by: guide.md §2.11 (generic expert labels give no accuracy benefit). [primary]
  194. Salewski, L., Alaniz, S., Rio-Torto, I., Schulz, E., and Akata, Z., "In-Context Impersonation Reveals Large Language Models' Strengths and Biases," NeurIPS 2023 (Spotlight). arXiv:2305.14930. https://arxiv.org/abs/2305.14930 Note: shows domain-specific, mixed role-prompting effects, not a flat negative. Cited by: guide.md §2.11 (role-prompting effects). [primary]
  195. Kong, A., Zhao, S., Chen, H., et al., "Better Zero-Shot Reasoning with Role-Play Prompting," NAACL 2024 (long papers). arXiv:2308.07702. https://arxiv.org/abs/2308.07702 / https://aclanthology.org/2024.naacl-long.228/ Note: role-play gains come from activating task-specific reasoning, not from credential labels. Cited by: guide.md §2.11 (process over credential). [primary]
  196. Schulhoff, S., et al., "The Prompt Report: A Systematic Survey of Prompting Techniques," arXiv:2406.06608, Jun 2024. https://arxiv.org/abs/2406.06608 Cited by: guide.md §2.8, §2.11 (broad prompt-structure and format-specification consensus). [primary]
  197. Anthropic, "Prompt engineering overview" (https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/overview) and "Be clear, direct, and detailed" (https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/be-clear-and-direct). Cited by: guide.md §2.9 (negative instructions can backfire), §2.11 (XML-tag structure). Practitioner guidance, not a peer-reviewed finding. [primary]
  198. Walters, W. H., and Wilder, E. I., "Fabrication and errors in the bibliographic citations generated by ChatGPT," Scientific Reports 13, 14045, Aug 2023. https://doi.org/10.1038/s41598-023-41032-5 Cited by: guide.md §2.10 (fabricated-citation rates). [primary]
  199. Haman, M., and Školník, M., "Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis," Journal of Medical Internet Research 26, e53164, 2024. https://www.jmir.org/2024/1/e53164 Cited by: guide.md §2.10 (fabricated-citation rates, secondary). [secondary]
  200. Rashkin, H., et al., "Measuring Attribution in Natural Language Generation Models" (AIS framework), Computational Linguistics 49(4), Dec 2023. arXiv:2112.12870. https://arxiv.org/abs/2112.12870 Companion: Gao, T., et al., "Enabling Large Language Models to Generate Text with Citations," EMNLP 2023, arXiv:2305.06983. Cited by: guide.md §2.10 (grounding answers in provided documents reduces hallucination). Note: the AIS paper defines the attribution metric; it does not itself report a 20–40% reduction figure, so the guide hedges the magnitude. [primary]
  201. Whiting v. City of Athens (6th Cir., March 2026), order imposing $15,000 per attorney plus opposing counsel's appellate fees and double costs for fabricated citations across three appellate briefs. Coverage: LawNext, "Sixth Circuit Slaps Steep Sanctions on Two Lawyers for Fake Citations," Mar 2026, https://www.lawnext.com/2026/03/sixth-circuit-slaps-steep-sanctions-on-two-lawyers-for-fake-citations-and-misrepresentations-in-appellate-briefs.html; Reason/Volokh, Mar 14, 2026. Added during legal audit, 2026-06-08; replaces an unverifiable "$31,000 C.D. Cal." figure. Cited by: guide.md §4.7. [primary order, secondary coverage]
  202. xAI, "Grok 4.3" model and pricing documentation, https://docs.x.ai/docs/models, beta Apr 17, 2026 / public API Apr 30–May 4, 2026. Confirms $1.25 per million input tokens, $2.50 per million output, 1M-token context, native video input, and reasoning on by default (defaults to reasoning_effort="low", always on). Added 2026-06-08 to source Grok 4.3-specific claims previously hung on [47], which covers the earlier Grok 4. Cited by: guide.md §1.2 / §1.3 / §2.5 / §6 / Appendix C. [time-sensitive]
  203. Vals AI, "Case Law v2" legal-reasoning leaderboard, https://www.vals.ai/benchmarks/case_law_v2, accessed 2026-06-08. Grok 4.3 ranked first at 79.31%. Added 2026-06-08. Note: this is a benchmark provider's own result, not an independent courtroom test. Cited by: guide.md §1.2 / §2.6. [time-sensitive]
  204. Law360 Pulse, "Finnegan Opens AI Practice With Dedicated IP Teams," https://www.law360.com/pulse/articles/2385794/finnegan-opens-ai-practice-with-dedicated-ip-teams, Sept 9, 2025. Finnegan formally launched "AI + Finnegan" in Sept 2025: four client-service teams (AI + Patent, AI + Copyright, AI + Privacy, AI + Trade Secrets), partner co-leads (Frank DeCosta, Karthik Kumar), and named client work including OpenAI. Managing partner James R. Barney quoted on the launch. Corroborated by Patent Lawyer Magazine (Sept 12, 2025), World IP Review (Sept 12, 2025), and Law.com / National Law Journal (Sept 9, 2025). Added 2026-06-08 during the §4.9 firm-accuracy pass. Cited by: guide.md §4.9. [primary press]
  205. Finnegan, "AI + KM" internal-program page (https://www.finnegan.com/en/firm/legal-value-solutions/ai-km.html) and the "AI + Finnegan" practice page (https://www.finnegan.com/en/work/practices/ai-finnegan/index.html), accessed 2026-06-08. Describes an internal AI Tech Tools Committee (partners, technologists, risk staff; partner co-lead Aaron Capron, independently confirmed in [204]), a Knowledge Management & Innovation Team (Director Benjamin Chi), prompt-engineering training for attorneys, and ISO 27001:2022 certification (2025). The firm states its tool bench mixes "proprietary and industry-standard tools" but names no specific vendor (no Copilot/Harvey/CoCounsel disclosure). Source is the firm's own materials, not third-party reporting. Added 2026-06-08. Cited by: guide.md §4.9. [firm self-disclosure]
  206. OpenRouter, X post describing Fusion model fanout and judge synthesis, https://x.com/OpenRouter/status/2065856862914785404, Jun 13, 2026. Local capture: raw/ai/openrouter-how-does-it-work-when-you-send-a-prompt-to-fusion-we-fan-it-out-to-a-.md. Describes sending one prompt to a panel of models in parallel, with tools enabled, then using a judge model to extract consensus, contradictions, partial coverage, and unique insights. Cited by: guide.md §1.4.3 / §2.10 / §5.6.1. [practitioner/vendor post; time-sensitive]
  207. Thariq, "Using Claude Code: The Unreasonable Effectiveness of HTML," X longform article, http://x.com/i/article/2052796100608974848, May 8, 2026. Local capture: raw/ai/using-claude-code-the-unreasonable-effectiveness-of-html-2052796100608974848.md. Argues that HTML artifacts can carry richer plans, diffs, diagrams, prototypes, and review surfaces than markdown for agent work. Cited by: guide.md §5.5 / Appendix B. [practitioner article]
  208. "Karpathy's 4 CLAUDE.md rules cut Claude mistakes from 41% to 11%. After 30 codebases, I added 8 more," X longform article, http://x.com/i/article/2053106718226227203, May 9, 2026. Local capture: raw/ai/karpathys-4-claude-md-rules-cut-claude-mistakes-from-41-to-11-after-30-codebases.md. Useful for practitioner rules around concise instruction files, surfacing conflicts, token budgets, checkpoints, read-before-write, and visible failure. Claims about exact mistake-rate reductions should be treated as self-reported unless independently verified. Cited by: guide.md §3.4 / Appendix B. [practitioner article; self-reported metrics]
  209. "How to Build Codex Knowledge Vault That Gets Smarter Every Day Without You Doing Anything," X longform article, http://x.com/i/article/2052866079391612933, May 9, 2026. Local capture: raw/ai/how-to-build-codex-knowledge-vault-that-gets-smarter-every-day-without-you-doing.md. Practitioner pattern for context debt, persistent knowledge layers, and vault topology around AGENTS.md, inbox, notes, ideas, and projects. Cited by: guide.md §3.4 / §5.5. [practitioner article]
  210. iFixAi GitHub repository, https://github.com/ifixai-ai/iFixAi, captured 2026-06-15; companion X longform article "Misalignment Is Unmeasured. iFixAi Is the Open-Source Diagnostic to Change That," http://x.com/i/article/2052027135619919876, May 6, 2026. Local captures: raw/ai/ifixai-ai-ifixai.md and raw/ai/misalignment-is-unmeasured-ifixai-is-the-open-source-diagnostic-to-change-that-2.md. Useful as an example of repeatable AI diagnostics, scorecards, cross-provider judging, and content-addressed manifests. Cited by: guide.md §3.8 / §5.9 / Appendix A. [open-source repo + practitioner article]
  211. CamoText, "Use Claude Cowork with Anonymization — Private Agentic AI Workflow," https://camotext.ai/blogposts/use-claude-cowork-with-anonymization, Mar 5, 2026. Local capture: raw/ai/use-claude-cowork-with-anonymization-private-agentic-ai-workflow.md. Describes local anonymization and metadata stripping before granting file-level access to an agentic AI service. Cited by: guide.md §4.6.2 / §5.8. [vendor article]
  212. Browse.sh, "open catalog of browser automation skills," https://browse.sh, captured 2026-06-15. Local capture: raw/ai/browse-sh-open-catalog-of-browser-automation-skills.md. Presents browser automation skills as reusable domain-specific actions with selectors, XHR/API knowledge, and cloud sessions. Cited by: guide.md §1.4.3 / §5.9 / Appendix A. [vendor product page; time-sensitive]
  213. Aakash Gupta, "Claude Skills: The 7 Laws from 75 Tests," https://www.aibyaakash.com/p/claude-skills-7-laws, captured 2026-06-15. Local capture: raw/ai/claude-skills-the-7-laws-from-75-tests-2026-guide.md. Practitioner guide on skills as reusable artifacts with descriptions, triggers, and scoped instructions. Cited by: guide.md §3.4 / Appendix B. [practitioner article; test claims self-reported]
  214. Boris Cherny, X post on moving from one agent to many agents, https://x.com/bcherny/status/2053982327123132846, May 11, 2026. Local capture: raw/ai/bcherny-the-best-way-to-level-up-from-1-agent-gt-many-agents-no-more-cycling-bet.md. Useful as a practitioner/product signal for multi-agent workspace flow. Cited by: guide.md §1.4.3 / §5.6. [practitioner post]
  215. Anthropic, claude-for-legal GitHub repository, https://github.com/anthropics/claude-for-legal, captured 2026-06-15. Local capture: raw/ai/anthropics-claude-for-legal.md. Official legal-workflow repository signal; useful for legal AI as packaged workflow artifacts rather than generic chat. Cited by: guide.md §4.5 / §4.10. [primary repo; time-sensitive]
  216. Anthropic, "Updates to Consumer Terms and Privacy Policy," https://www.anthropic.com/news/updates-to-our-consumer-terms, Aug 28, 2025. Consumer Claude users (Free, Pro, Max) must choose whether to allow their chats to be used for model training; those who allow it have data retained up to five years, those who decline keep the prior 30-day window. Business/API tiers are unaffected. Cited by: guide.md §2.7. [primary; vendor]
  217. Anthropic, "Claude Cowork" (Anthropic Labs), https://www.anthropic.com/news/claude-cowork, Jan 12, 2026. Desktop agent that works across a user's files and documents; a consumer agent surface rather than a repo-and-terminal coding tool. Cited by: guide.md §2.12 / §6.2. [primary; vendor; time-sensitive]
  218. Reporting on Apple–Google Gemini deal for Siri, e.g. TechCrunch, "Apple to pay Google to power a revamped Siri with custom Gemini model," https://techcrunch.com/2026/01/12/apple-google-gemini-siri/, and CNBC coverage, Jan 12, 2026. Custom Gemini models to power a more personalized Siri later in 2026. Cited by: guide.md §2.4. [secondary news; time-sensitive]
  219. Meta, "Llama Community License Agreement," https://www.llama.com/llama3/license/. Entities with greater than 700 million monthly active users must request a separate license rather than relying on the community license. Cited by: guide.md §1.2. [primary; vendor]
  220. Google DeepMind, "SynthID," https://deepmind.google/technologies/synthid/. Imperceptible watermark embedded in AI-generated images (including Gemini/Nano Banana output) for later identification. Cited by: guide.md §2.3. [primary; vendor]
  221. A&O Shearman / Microsoft / Harvey, "ContractMatrix," announced Dec 21, 2023, https://www.allenovery.com/en-gb/global/news-and-insights/news/ao-announces-exclusive-launch-partnership-with-harvey. Co-built contract-analysis tool. Cited by: guide.md §4.4. [primary; firm announcement]
  222. Latham & Watkins firmwide Harvey rollout, Aug 11, 2025, https://www.lw.com/en/news/latham-watkins-expands-firmwide-access-to-harvey-ai. Deployed to more than 3,600 attorneys. Cited by: guide.md §4.4. [primary; firm announcement]
  223. Anthropic, "Models overview," https://platform.claude.com/docs/en/docs/about-claude/models/overview. Claude Fable 5 general availability June 9, 2026; described as Anthropic's most capable widely released model, $10 per million input tokens and $50 per million output tokens, 1M-token context window, 128k max output. Cited by: guide.md §1.2, §1.3, §6.1, Appendix C, Appendix D. [primary; vendor]
  224. Google, "Gemini API pricing," https://ai.google.dev/pricing, and "Gemini models," https://ai.google.dev/gemini-api/docs/models. Gemini 3.5 Flash stable flagship pricing $1.50 per million input tokens and $9 per million output tokens; Gemini 3.1 Pro remains in preview. Cited by: guide.md §1.2, §1.3, §6.1, Appendix C, Appendix D. [primary; vendor]
  225. OpenAI, "Introducing the Codex app," https://openai.com/index/introducing-the-codex-app/. Desktop application for supervising multiple Codex agents; Windows support added Mar 4, 2026; distinct from the Codex CLI; reported 2M+ weekly active users. Cited by: guide.md §3.6, §6.2. [primary; vendor]
  226. OpenAI, "Introducing workspace agents in ChatGPT," https://openai.com/index/introducing-workspace-agents-in-chatgpt/. Codex-powered shared agents with scheduled runs, approval gates, and Agent Analytics; available on Business, Enterprise, and Edu; credit-based pricing from May 6, 2026. Cited by: guide.md §2.12, §3.3, §3.6. [primary; vendor]
  227. OpenAI, "Introducing apps in ChatGPT and the Apps SDK," https://openai.com/index/introducing-apps-in-chatgpt/. Developer preview Nov 13, 2025; lets developers build interactive apps that run inside ChatGPT. Cited by: guide.md §3.6. [primary; vendor]
  228. OpenAI, "Introducing Lockdown Mode and Elevated Risk labels in ChatGPT," https://openai.com/index/introducing-lockdown-mode-and-elevated-risk-labels-in-chatgpt/. Rolling out June 4, 2026; Lockdown Mode disables web browsing, image generation, Deep Research, Agent Mode, connectors, and downloads to reduce prompt-injection exposure. Cited by: guide.md §2.7, §5.8. [primary; vendor]
  229. Google, "URL context," https://ai.google.dev/gemini-api/docs/url-context (Gemini API primitive that fetches and grounds on URLs), and OpenAI, "Built-in tools," https://developers.openai.com/api/docs/guides/tools (Responses API web_search and computer-use tools). Cited by: guide.md §2.5, §5.9. [primary; vendor]
  230. Anthropic, "Recursive self-improvement," https://www.anthropic.com/institute/recursive-self-improvement. "As of May 2026, more than 80% of the code we merge into Anthropic's codebase was authored by Claude." Cited by: guide.md §5.1. [primary; vendor]
  231. Thomson Reuters, "Thomson Reuters launches CoCounsel Legal MCP," press release May 12, 2026, https://www.thomsonreuters.com/en/press-releases/2026/may/thomson-reuters-launches-cocounsel-legal-mcp.html. Connects Claude and other MCP clients to CoCounsel Legal and Westlaw's 1.9B documents. Cited by: guide.md §4.5. [primary; vendor press release]
  232. Thomson Reuters, "Introducing HighQ MCP," May 19, 2026, https://legal.thomsonreuters.com/blog/introducing-highq-mcp/. Read-only, permission-controlled, fully audit-logged MCP server for HighQ matter data. Cited by: guide.md §4.6.1. [primary; vendor]
  233. Anthropic, "Claude for Financial Services," https://www.anthropic.com/news/claude-for-financial-services, and "Claude for Life Sciences / Healthcare" coverage, Jan 2026; HIPAA-ready configurations for regulated workflows. Cited by: guide.md §6.1. [primary; vendor]
  234. Anthropic, "About Claude's usage limits," https://support.claude.com/en/articles/9797557, with https://support.claude.com/en/articles/8325606 and https://support.claude.com/en/articles/11145838. Usage is metered in a rolling 5-hour session window plus separate weekly limits; Claude Code draws from the same pool as chat. Cited by: guide.md §2.13. [primary; vendor support]
  235. Anthropic, "Learn," https://www.anthropic.com/learn. Learning hub linking Anthropic Academy courses (hosted on Skilljar) for users and builders. Cited by: guide.md §2.12, §2.13, Appendix D. [primary; vendor]
  236. Google, "I/O 2026 developer highlights: Antigravity, Gemini API, AI Studio," https://blog.google/innovation-and-ai/technology/developers-tools/google-io-2026-developer-highlights/, May 2026, and "Introducing Managed Agents in the Gemini API," https://blog.google/innovation-and-ai/technology/developers-tools/managed-agents-gemini-api/, May 19, 2026. Managed Agents use the Antigravity agent harness, run in isolated Linux environments, and can be defined with AGENTS.md and SKILL.md files. Cited by: guide.md §1.2, §1.4.3, §5.9, §6.2. [primary; vendor; time-sensitive]
  237. Anthropic, "Extend Claude with skills," https://docs.anthropic.com/en/docs/claude-code/skills, and "Skill authoring best practices," https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices. Skills are discoverable instruction bundles with scripts and resources; good skills are concise, structured, and tested against real usage. Cited by: guide.md §3.4, §5.6, §5.9, Appendix D. [primary; vendor]
  238. OpenAI, "ChatGPT Enterprise & Edu release notes," https://help.openai.com/en/articles/10128477-chatgpt-enterprise-edu-release-notes. May 2026 release notes describe workspace agents becoming generally available in ChatGPT Business, Enterprise, and Edu, plus admin controls, app-specific action safeguards, Codex goal mode, browser improvements, remote locked use, analytics, and plugin-sharing status. Cited by: guide.md §1.4.3, §3.6, §6.2. [primary; vendor; time-sensitive]
  239. Anthropic, "Introducing Claude Opus 5," https://www.anthropic.com/news/claude-opus-5, July 24, 2026. Released July 24, 2026 at $5/$25 per MTok, 1M context, 128K max output, thinking on by default, adjustable effort (low through max). Anthropic claims performance within 0.5% of Claude Fable 5's peak on CursorBench 3.2 at maximum effort at half the cost per task, and more than double Opus 4.8 on Frontier-Bench v0.1. Benchmarks are the lab's own. Cited by: guide.md §1.2, §1.3, §3.5, §6.1, Appendix C. [primary; vendor; time-sensitive]
  240. Anthropic, "Pricing," https://platform.claude.com/docs/en/about-claude/pricing, accessed August 31, 2026. Current model pricing per MTok: Fable 5 and Mythos 5 $10/$50, Opus 5 / 4.8 / 4.7 / 4.6 / 4.5 $5/$25, Sonnet 5 $2/$10, Sonnet 4.6 / 4.5 $3/$15, Haiku 4.5 $1/$5. Confirms that Sonnet 5's $2/$10 introductory rate became standard and the scheduled September 1, 2026 increase to $3/$15 was cancelled. Also documents cache-read pricing at 0.1x input and the 50% Batch API discount. Cited by: guide.md §1.2, §1.3, §3.5, §4.3.1, §6.1, Appendix C. [primary; vendor; time-sensitive]
  241. OpenAI, "Pricing," https://developers.openai.com/api/docs/pricing, and "GPT-5.6 Sol," https://developers.openai.com/api/docs/models/gpt-5.6-sol, accessed August 31, 2026. Current rates per MTok are Sol $4/$20, Terra $2/$12, and Luna $0.20/$1.20. OpenAI says Sol's promotional rate will remain available at least through November 21, 2026, but announces no later price. Requests above 272,000 input tokens incur 2x input and 1.5x output pricing. Cache reads can cost 0.1x standard input and cache writes can cost 1.25x. Batch availability and discounts depend on the model and endpoint. Cited by: guide.md §1.2, §1.3, §3.5, §4.3.1, §6.1, Appendix C. [primary; vendor; time-sensitive]
  242. Google, "Gemini API pricing," https://ai.google.dev/gemini-api/docs/pricing, accessed August 31, 2026, and "Introducing Gemini 3.7 Flash," https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/. Gemini 3.7 Flash and 3.6 Flash at $0.75/$3.75 per MTok on introductory pricing through December 31, 2026, rising to $1.50/$7.50 on January 1, 2027. Gemini 3.1 Pro Preview $2/$12; Gemini 3.5 Flash $1.50/$9; 3.5 Flash-Lite $0.30/$2.50. Google describes 3.7 Flash as its "most intelligent workhorse model," not as its flagship; the Pro line remains nominally top-of-range. Gemini 4 confirmed in pre-training with no announced date, specs, or pricing. Cited by: guide.md §1.2, §1.3, §3.5, §6.1, Appendix C. [primary; vendor; time-sensitive]
  243. xAI, "Introducing Grok 4.6," https://x.ai/news/grok-4-6, August 12, 2026. Released August 12, 2026 at $2/$6 per MTok, with a faster variant at 2x. xAI positions it for long-running agents and interactive or visual work. Context window reported at 500,000 tokens, the smallest among current frontier models, which otherwise sit at or above 1M. No published Vals AI Case Law result covers Grok 4.6, so the Grok 4.3 legal-reasoning ranking in [203] should not be assumed to carry forward. Cited by: guide.md §1.2, §1.3, §6.1, Appendix C. [primary; vendor; time-sensitive]
  244. Meta AI Research, "Introducing Muse Glimmer," https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model, and https://developer.meta.com/ai/models/muse-glimmer/, August 10, 2026. Muse Glimmer is a 30B-parameter agent-optimized model under an unmodified Apache 2.0 license, materially more permissive than the Llama community license. Muse Spark 1.2 (reported August 5, 2026, 1M context, coding-focused) remains closed-weight: Zuckerberg said weights are "coming soon" without naming a date or license as of August 11, 2026. Cited by: guide.md §1.2, §6.1. [primary for Glimmer; secondary and unconfirmed for Spark 1.2 open-weighting]
  245. Anthropic, "Introducing dynamic workflows in Claude Code," https://claude.com/blog/introducing-dynamic-workflows-in-claude-code, announced May 28, 2026, subsequently marked generally available. Claude writes orchestration scripts that run tens to hundreds of parallel sub-agents in one session, including agents tasked with refuting other agents' findings, iterating until answers converge, with automatic progress saving. Anthropic's own documentation warns the feature consumes substantially more tokens than a typical session. Cited by: guide.md §5.6, §6.2. [primary; vendor]
  246. Reporting on Anthropic's self-hosted environments public beta for Claude Code, opened August 6, 2026: https://www.unite.ai/claude-code-sessions-can-now-run-on-infrastructure-your-team-controls/ and https://enterprisedna.co/resources/news/anthropic-claude-code-self-hosted-runner-enterprise-2026/. A claude self-hosted-runner command turns customer-controlled machines and containers into the compute layer for Claude Code sessions, available to Team and Enterprise plans, off by default. NOT CONFIRMED against Anthropic's own documentation. Verify before relying on it for a compliance commitment. Cited by: guide.md §5.7, §6.2. [secondary; unverified; time-sensitive]
  247. OpenAI, "OpenAI launches the Deployment Company to help businesses build around intelligence," https://openai.com/index/openai-launches-the-deployment-company/, May 11, 2026. A professional services business backed by more than $4 billion from nineteen investment firms, consultancies, and systems integrators, majority-owned and controlled by OpenAI, embedding "forward deployed engineers" into customer organizations. Launched with roughly 150 such engineers via the acquisition of applied-AI consultancy Tomoro. Co-lead investors include TPG, Advent, Bain Capital, and Brookfield. Cited by: guide.md §1.2. [primary; vendor]
  248. Anthropic, "Claude for the legal industry," https://claude.com/blog/claude-for-the-legal-industry, May 12, 2026; repository at https://github.com/anthropics/claude-for-legal; coverage at https://www.artificiallawyer.com/2026/05/12/claude-for-legal-launches-may-reshape-the-legal-tech-world/ and https://www.lawnext.com/2026/05/anthropic-goes-all-in-on-legal-releasing-more-than-20-connectors-and-12-practice-area-plugins-for-claude.html. Open-source suite of 12 practice-area plugins (commercial, employment, privacy, product, corporate, AI governance, litigation associate, law student, and others), a large set of named workflow agents, and more than 20 MCP connectors including DocuSign, Ironclad, iManage, NetDocuments, LexisNexis, Thomson Reuters CoCounsel, Box, Everlaw, Relativity, and Datasite. Available to paid Claude customers with enterprise admin control over enablement. Freshfields, Quinn Emanuel Urquhart & Sullivan, Holland & Knight, and Crosby Legal named as using Claude on live matters. Cited by: guide.md §4.10.1, §6.1, §6.3. [primary; vendor; corroborated by trade press]
  249. Legora, "Legora raises $550 million Series D to fuel US growth," https://legora.com/newsroom/legora-raises-550-million-series-d-to-fuel-us-growth, March 2026, and https://news.crunchbase.com/venture/unicorn-legal-tech-ai-startup-legora-triples-valuation/. $550M Series D at a $5.55B valuation led by Accel. Named customers include Cleary Gottlieb, White & Case, Linklaters, Goodwin, Bird & Bird, and Dentons. Separately, https://www.cityam.com/legora-eyes-10bn-funding-valuation-four-months-after-last-raise/ and https://techstartups.com/2026/08/13/legal-ai-startup-legora-seeks-10-billion-valuation-nearly-doubling-its-worth-in-just-four-months/ report August 2026 early-stage talks at a valuation above $10B. The $10B figure is REPORTED, not closed. Cited by: guide.md §4.2. [primary for Series D; secondary and unconfirmed for the $10B round]
  250. Kirkland & Ellis, "Kirkland & Ellis and Palantir Partner to Transform Private Equity Fundraising with Exclusive AI-Powered Fund Enterprise Platform," https://www.kirkland.com/news/press-release/2026/06/kirkland-ellis-and-palantir-p-to-transform-pe-fundraising-with-excl-ai-pow-fund-enterprise-plat, June 4, 2026. Multiyear partnership building an exclusive, proprietary platform covering fund documentation, investor solutions, side letter drafting, obligation tracking, closing commitments, and ongoing compliance. Kirkland's Investment Funds Group comprises more than 1,000 lawyers. The firm supported nearly $500 billion in capital raised or targeted for clients in 2025. Exclusive to Kirkland, therefore a competitive-differentiation signal rather than a procurement option for other firms. Cited by: guide.md §4.4. [primary; firm press release]
  251. Norton Rose Fulbright, "AI in litigation: Update on Gen AI sanctions in 2026," https://www.nortonrosefulbright.com/en-us/knowledge/publications/792d8bf3/ai-in-litigation-update-on-gen-ai-sanctions-in-2026, June 2026. Documents more than 1,148 US attorney-hallucination cases tracked as of publication. Case table: Fletcher v. Experian (5th Cir., Feb. 18, 2026, $2,500); In re: Nwaubani (4th Cir., Mar. 11, 2026, public admonishment); Whiting v. City of Athens (6th Cir., Mar. 13, 2026, $15,000 each plus fees, double costs, disciplinary referral); United States v. Farris (6th Cir., 2026, case removal and disciplinary referral); Gamez v. County of Fresno (E.D. Cal., Apr. 9, 2026, no sanctions); Fivehouse v. DOD (E.D.N.C., Apr. 27, 2026, public reprimand). Concludes no Gen-AI-specific rule proved necessary: existing FRAP, FRCP, and state professional conduct rules sufficed. Cited by: guide.md §4.7. [secondary; law firm publication; well-sourced]
  252. California Senate Bill 574 (2025–2026), official text and status, https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260SB574 and https://leginfo.legislature.ca.gov/faces/billStatusClient.xhtml?bill_id=202520260SB574, accessed August 31, 2026. The August 21 Assembly amendment would prohibit delegating the practice of law to generative AI, require reasonable verification and correction of AI output, require disclosure of AI use for documents submitted to court, and prohibit any filed citation that the responsible attorney has not personally verified. It also restricts nonpublic inputs and arbitrator delegation. Status on August 31: active bill in Assembly floor process, ordered to third reading, not law. Cited by: guide.md §4.6. [primary; legislature; time-sensitive]
  253. Consumer AI subscription pricing as of August 31, 2026. Google One official US plans, https://one.google.com/about/plans, lists Google AI Plus at $9.99 per month and AI Pro at $19.99. Other plan figures in the guide remain drawn from vendor pages where available and aggregators where a stable official page was not captured. Confirm geography, billing cadence, and current terms before budgeting. Cited by: guide.md §1.2, §1.3, Appendix C. [mixed primary and secondary; time-sensitive]
  254. AI prompt discovery and work product in 2026. Conservation Law Foundation, Inc. v. Shell Oil Co., docket at https://www.courtlistener.com/docket/60042265/conservation-law-foundation-inc-v-shell-oil-company/, including ECF 970, ordered production of a testifying expert's AI prompts concerning document-culling methodology and stayed production pending Rule 72(a) review. Morgan v. V2X, Inc., No. 1:25-cv-01991, ECF 65 (D. Colo. Mar. 30, 2026), https://law.justia.com/cases/federal/district-courts/colorado/codce/1:2025cv01991/245077/65/, held that Rule 26(b)(3) could protect a pro se litigant's AI-assisted litigation preparation, while requiring disclosure of the tool and adding contractual safeguards to a protective order. Assini v. Hayward, 2026 NY Slip Op 26086, https://www.nycourts.gov/reporter/current/3dseries/2026/2026_26086.shtml, adopted similar reasoning and quashed a subpoena to OpenAI. Cited by: guide.md §4.6.3. [primary dockets and opinions; context-specific]
  255. Thomson Reuters, "Thomson Reuters Launches Next Generation of CoCounsel Legal," https://www.thomsonreuters.com/en/press-releases/2026/august/thomson-reuters-launches-next-generation-of-cocounsel-legal-the-ai-ecosystem-built-for-legal-professionals, August 20, 2026. Includes Westlaw Brief Builder, Deep Research Verify (checks cited propositions against underlying sources), and Tabular Analysis (up to 100 structured questions across up to 10,000 documents). Vendor accuracy claims are unaudited. The Stanford RegLab study (May 2024) remains the only preregistered public benchmark of citation-grounded legal tools and found hallucination rates of 17% to 33%; no independent 2026 re-test has published. Cited by: guide.md §4.11, §6.1. [primary vendor announcement; accuracy claims unverified]
  256. E-discovery and litigation-support generative AI shift, 2026. Relativity made aiR for Review and aiR for Privilege standard in RelativityOne and Everlaw made single-use AI tasks free: https://www.ilsteam.com/relativity-and-everlaws-landmark-free-gen-ai-announcements-level-the-playing-field/. DISCO launched Cecilia agentic AI for multi-step fact investigation, February 2026: https://www.lawnext.com/2026/02/disco-launches-scaled-agentic-ai-tool-for-large-discovery-and-fact-investigation-matters.html. Chronology tooling: https://www.prweb.com/releases/casefleets-agentic-ai-builds-litigation-ready-case-chronologies-in-minutes-302805671.html. Medical-record review: EvenUp Piai / MedChrons, vendor-sourced, https://www.evenuplaw.com/guides/ai-medical-records-summary-for-lawyers/. Vendor cost-reduction claims (40% to 60%) are unaudited. Cited by: guide.md §4.11. [mixed primary vendor and secondary trade press; accuracy and savings claims unverified]
  257. Trellis MCP connectors. "Trellis Brings the Largest State Trial Court Dataset in the U.S. to Claude," https://www.prnewswire.com/news-releases/trellis-brings-the-largest-state-trial-court-dataset-in-the-us-to-claude--and-anyone-can-try-it-for-free-302768874.html, May 12, 2026; "Trellis Expands Agentic Access to the Nation's Largest State Trial Court Dataset Through ChatGPT and Trellis Chat," https://www.prnewswire.com/news-releases/trellis-expands-agentic-access-to-the-nations-largest-state-trial-court-dataset-through-chatgpt-and-trellis-chat-building-on-its-claude-mcp-connector-302836491.html, July 28, 2026. 2.5B+ state trial court records reachable from Claude and ChatGPT. Judge and motion analytics are base-rate statistics, not case-specific forecasts. Cited by: guide.md §4.11, §6.3. [primary vendor announcement]
  258. Oregon AI-citation sanctions, primary orders. Couvrette v. Wisnovsky, No. 1:21-cv-00157, sanctions order entered December 12, 2025 ($15,500), https://websitedc.s3.amazonaws.com/documents/Couvrette_v._Wisnoksy_USA_12_December_2025.pdf, and fee-and-cost order entered March 23, 2026 ($94,704.38), https://docs.justia.com/cases/federal/district-courts/oregon/ordce/1%3A2021cv00157/158388/225. Doiban v. Oregon Liquor & Cannabis Commission, 347 Or. App. 742, 747–48 (2026), official opinion at https://ojd.contentdm.oclc.org/digital/api/collection/p17027coll5/id/41541/download, imposed $10,000 for fifteen false citations and nine contrived quotations. William L. Ghiorso was counsel. In re Ghiorso is a separate 2016 disciplinary matter. The reported $145,000 Q1 aggregate could not be reproduced and is omitted. Cited by: guide.md §4.7. [primary orders and opinion]
  259. Model Context Protocol, specification revision 2026-07-28, https://blog.modelcontextprotocol.io/posts/2026-07-28/, July 28, 2026. Rewrites MCP as a stateless request/response protocol rather than long-lived bidirectional sessions, hardens OAuth with issuer validation per RFC 9207 and audience-bound tokens, requires application_type at registration, and deprecates Dynamic Client Registration in favor of Client ID Metadata Documents. Directly targets the ambient-authority risk described in [111]. Cited by: guide.md §5.8. [primary; standards body]
  260. "Exposed by Design: A Large-Scale Security Assessment of Internet-Facing MCP Servers," arXiv:2608.00150, https://arxiv.org/abs/2608.00150, submitted July 31, 2026. More than 21,000 discoverable MCP servers, 640 confirmed production servers. Of those dynamically audited, 91.8% lacked OAuth entirely and 687 tool instances exposed raw shell access. First large-scale dynamic (not merely static) assessment of the MCP ecosystem. Updates and strengthens the OX Security figures in [100][137]. Cited by: guide.md §5.8. [primary; preprint, not yet peer reviewed]
  261. "Agent Data Injection," arXiv:2607.05120, https://arxiv.org/abs/2607.05120, submitted July 6, 2026. An indirect-injection class using probabilistic delimiter injection: fake punctuation and delimiter characters planted in structured data fields corrupt what an agent trusts as legitimate metadata, rather than planting instruction-shaped text. Demonstrated against Claude Code, Claude in Chrome, OpenAI Codex, Gemini CLI, and Google Antigravity in default configurations. Not addressed by instruction-pattern detection; the mitigation category is trusted/untrusted data-boundary sanitization. Cited by: guide.md §5.8. [primary; preprint, not yet peer reviewed]
  262. Anthropic, "Inference hooks," https://platform.claude.com/docs/en/manage-claude/inference-hooks, beta, August 5, 2026. Holds governed prompts across Claude, Cowork, and Claude Code for an allow/deny verdict from the organization's own AI security server before inference proceeds, with denials logged to a compliance activity feed. Structural counterpart to OpenAI's Lockdown Mode [228]. Cited by: guide.md §5.8. [primary; vendor]
  263. Remote and mobile agent session control, summer 2026. Claude Cowork reached web, iOS, and Android July 7, 2026: https://techcrunch.com/2026/07/07/the-coding-agent-wars-are-spilling-into-the-rest-of-the-office-claude-cowork/. OpenAI Codex Remote reached general availability across ChatGPT plans in late June 2026: https://developers.openai.com/codex/changelog (exact date corroborated by dated press rather than read directly from the changelog entry). Google Antigravity added Remote Control in v2.9.1 on August 20, 2026: https://antigravity.google/changelog?tab=hub. Sessions continue running on the machine holding local files and credentials while being steered from another device. Cited by: guide.md §6.2. [mixed primary changelog and secondary press; Codex date is secondary]
  264. Moonshot AI, Kimi K3, https://huggingface.co/moonshotai/Kimi-K3, weights published July 27, 2026. 2.8T total parameters / 104B activated, mixture-of-experts, 1M context, released under a custom "Kimi K3 License" rather than Apache 2.0 or an OSI-approved license. Read the license before assuming commercial self-hosting rights. Cited by: guide.md §6.1, §5.7. [primary; model card]
  265. OpenAI, "Introducing ChatGPT Health," https://openai.com/index/introducing-chatgpt-health/, January 7, 2026; July 23 expansion reporting at https://techcrunch.com/2026/07/23/openai-makes-chatgpt-health-available-to-all-u-s-users/. The expansion covered US users aged 18 and older on web and iOS, subject to rollout limits. HHS mobile-health privacy guidance, https://www.hhs.gov/hipaa/for-professionals/privacy/guidance/cell-phone-hipaa/index.html, explains that consumer apps usually fall outside HIPAA unless acting for a covered entity or business associate. On the multistate attorney general subpoena: https://therecord.media/chatgpt-health-draws-concern-privacy-critics. Cited by: guide.md §2.14. [primary for product and HIPAA framework; secondary for rollout and enforcement reporting]
  266. CBS News, reporting on a July 23, 2026 lawsuit filed against OpenAI in Florida alleging that ChatGPT identified the plaintiff as having dysautonomia from uploaded labs and imaging and advised remaining recliner-bound, followed weeks later by a pulmonary embolism a physician attributed to the resulting immobility, https://www.cbsnews.com/news/chatgpt-dangerous-medical-advice-openai-lawsuit/, July 2026. ALLEGATIONS ARE UNPROVEN. Cited in guide.md for the behavioral pattern (confident diagnostic output without appropriate refusal), not for any finding of liability. Cited by: guide.md §2.14. [secondary; pending litigation; allegations unproven]
  267. PYMNTS, "As Tax Deadline Approaches, Consumers Are Going to AI Before Filing," https://www.pymnts.com/taxes/2026/as-tax-deadline-approaches-consumers-are-going-to-ai-before-filing/, 2026. A test ran eight fictional tax scenarios through ChatGPT, Gemini, Claude, and Grok with required forms supplied. Reporting described an average absolute error above $2,000 across the four chatbots and eight scenarios, with every model making at least one error. It did not establish that each model separately averaged more than $2,000. Cited by: guide.md §2.14. [secondary; trade publication reporting a controlled test]
  268. Consumer booking and purchasing agents, 2026. Google Gemini Agent multi-step tasks: https://support.google.com/gemini/answer/16596215. Chrome automatic browsing background: https://www.techspot.com/news/112334-project-mariner-dead-but-google-browser-controlling-ai.html. Delta Concierge autonomous rebooking and a documented failure in which the assistant told a passenger her flight was safe, then cancelled it and quoted double to rebook: https://www.pymnts.com/earnings/2026/delta-ai-assistant-boosts-customer-satisfaction-scores-25-points-during-travel-disruptions/ and https://viewfromthewing.com/deltas-ai-told-a-passenger-her-flight-home-was-safe-then-canceled-it-and-demanded-double-to-rebook/. ACI Worldwide consumer survey, June 28, 2026: 60% of UK consumers would stop using an AI shopping agent after one mistake and only 19% trust one to make everyday purchase decisions correctly, https://investor.aciworldwide.com/news-releases/news-release-details/six-ten-uk-consumers-would-stop-using-ai-shopping-agent-after. Cited by: guide.md §2.14. [mixed primary vendor documentation and secondary reporting]
  269. Journal of Accountancy, "Elder fraud rises as scammers use AI," https://www.journalofaccountancy.com/issues/2026/apr/elder-fraud-rises-as-scammers-use-ai/, April 2026. Voice cloning from short social-media audio samples used to stage fake family emergencies. Recommended countermeasure is a family safe word verified by callback to a known number. Cited by: guide.md §2.14. [secondary; professional association publication]
  270. Gartner survey of 782 IT infrastructure and operations leaders conducted November to December 2025, reported in The Register, https://www.theregister.com/2026/04/07/ai_returns_gartner/, April 7, 2026. 28% of AI infrastructure projects delivered promised ROI, 20% failed outright, 57% of organizations reported at least one AI failure. Leading causes were skill gaps and poor data quality, each cited by 38%. Cited by: guide.md §3.10. [secondary reporting of a primary survey]
  271. Boston Consulting Group, "CEOs report cost and revenue benefits from AI but are struggling to scale," https://www.bcg.com/press/22july2026-ceos-cost-revenue-benefits-ai-struggling-scale, July 22, 2026. Nearly nine in ten CEOs report some cost or revenue benefit in targeted areas while most struggle to scale AI enterprise-wide. Cited by: guide.md §3.10. [primary; consultancy press release; note the publisher sells AI advisory services]
  272. Microsoft Copilot enterprise adoption analysis, https://valueaddvc.com/blog/microsoft-copilot-enterprise-adoption-what-the-data-shows-about-real-usage-vs-hype, 2026. Aggregated governance-vendor and market data placing weekly active use of paid Copilot seats at roughly 20% to 30%, with workplace conversion around 35.8% versus roughly 83% for ChatGPT among users with access. Aggregated secondary data rather than a Microsoft disclosure. Cited by: guide.md §3.10. [secondary; aggregator]
  273. Forrester, Predictions 2026 (Future of Work), summarized in Computerworld, https://www.computerworld.com/article/4084372/analysts-companies-will-face-setbacks-after-ai-layoffs.html, and ITPro, https://www.itpro.com/business/business-strategy/analysts-warn-ai-layoffs-could-spark-a-new-wave-of-offshoring. Reporting attributes the 55% AI-layoff regret figure to Forrester. Figures about rehiring 35.6% of workers and spending more to restaff appear to come from a separate Careerminds survey of 600 HR professionals; its primary report was not obtained, so the guide omits them. Cited by: guide.md §3.10. [secondary reporting; attribution corrected]
  274. Bloomberg, "Commonwealth Bank reverses job cuts decision over AI chatbots," https://www.bloomberg.com/news/articles/2025-08-21/commonwealth-bank-reverses-job-cuts-decision-over-ai-chatbots, August 21, 2025. CBA cut 45 call-center roles citing a voice bot that had purportedly reduced call volume by 2,000 per week; call volumes then rose, the bank apologized for the error, and moved to rehire and redeploy staff. Predates the 2026 window but remains the best-verified named-company reversal and is used here for the pattern of a self-reported metric being the thing that was wrong. Cited by: guide.md §3.10. [primary reporting; 2025]
  275. Cloud Security Alliance, "New Cloud Security Alliance Survey Reveals 82% of Enterprises Have Unknown AI Agents in Their Environments," https://cloudsecurityalliance.org/press-releases/2026/04/21/new-cloud-security-alliance-survey-reveals-82-of-enterprises-have-unknown-ai-agents-in-their-environments, April 21, 2026. The January 2026 online survey covered 418 IT and security professionals. Token Security financed the work and co-developed the questionnaire. Sixty-five percent reported an AI-agent-related incident during the prior year, not an incident proven to have been caused by an agent; 61% reported data exposure. Cited by: guide.md §3.10. [primary survey release; vendor-commissioned]
  276. Microsoft Security Response Center, CVE-2025-32711, https://msrc.microsoft.com/update-guide/vulnerability/CVE-2025-32711, and Aim Security, "EchoLeak," with research paper https://arxiv.org/abs/2509.10540. Researchers demonstrated a zero-click attack path in which a later user query retrieved a crafted email and triggered hidden instructions capable of exposing data available through connected Microsoft services. Microsoft patched the issue before public disclosure and reported no known customer impact. Cited by: guide.md §3.10. [primary vendor advisory and research paper]
  277. Microsoft, Agent 365, https://www.microsoft.com/en-us/microsoft-agent-365, launched May 1, 2026 at $15 per user per month as a control plane placing AI agents under unified identity, permissions, and audit; June 2026 update expanded Purview Audit coverage to locally built and developer-built agents, https://techcommunity.microsoft.com/blog/agent-365-blog/whats-new-in-agent-365-%E2%80%93-june-2026/4535107. Requires the Microsoft 365 identity stack, so realistic mainly for organizations already standardized on it. Cited by: guide.md Appendix C. [primary; vendor]
  278. Bunce v. Visual Technology Innovations, No. 2:23-cv-01740 (E.D. Pa., Judge Kai N. Scott). Rule 11 AI-sanctions memorandum and order entered April 20, 2026 at docket entries 225–226, imposing $5,000 plus mandatory AI-ethics CLE on attorney Mathu G. Rajan. Docket: https://www.courtlistener.com/docket/67332015/. Contemporaneous account: https://reason.com/volokh/2026/04/20/it-will-be-your-name-and-license-on-the-line-not-chatgpts/. CORRECTION RECORD: guide.md v0.11.2 and earlier cited this matter as "2025 WL 4231632 (E.D. Pa. Jan. 21, 2025)," which is wrong on both the reporter number and the date. The error entered via source [96] (Bundle B) and was caught by adversarial docket review on 2026-08-31. Cited by: guide.md §4.7. [primary; docket]
  279. Harvey current scale and reported August 2026 round. Company figures at https://www.harvey.ai/ and https://www.harvey.ai/newsroom, now describing roughly 200,000 lawyers and 2,400+ organizations, substantially above the March 2026 figures in [67]. Reported August 7, 2026 funding talks at approximately $15.5B valuation on $350M+ ARR, and the July 2026 acquisition of Benchmark plus a Datasite partnership, via https://www.law360.com/pulse/amp/articles/2510984 and https://sacra.com/c/harvey/. The $15.5B figure is REPORTED, not closed, and is included alongside the Legora $10B report in [249] so the comparison is not one-sided. Cited by: guide.md §4.2. [primary vendor figures; secondary and unconfirmed for the round]
  280. ABA Formal Opinion 512 fee analysis, discussed in ABA Law Practice Today, "The Reasonableness of Fees When Using AI," https://www.americanbar.org/groups/law_practice/resources/law-practice-today/2024/december-2024/the-reasonableness-of-fees-when-using-ai/, December 2024. Under Rule 1.5 as construed by Opinion 512 [83]: an hourly biller may bill only time actually spent and may not bill for time saved by generative AI; may not generally bill clients for time spent learning the tool; must disclose in advance whether GAI costs are passed through; and a flat fee left unchanged after AI compresses the work may itself become unreasonable. NOTE: americanbar.org returned 403 to automated fetch during this pass, so these are reported summaries. Read the fee section of Opinion 512 directly before setting firm policy. Cited by: guide.md §4.6. [secondary; ABA publication summarizing a primary ABA opinion; not directly fetched]
  281. Privilege, work product, and AI use. United States v. Heppner, No. 1:25-cr-00503-JSR, ECF 27, 2026 WL 436479 (S.D.N.Y. Feb. 17, 2026), https://storage.courtlistener.com/recap/gov.uscourts.nysd.652137/gov.uscourts.nysd.652137.27.0.pdf, held on consumer-Claude facts that material created by a represented defendant without counsel's direction was neither privileged nor work product and that disclosure under the consumer service's terms waived otherwise applicable privilege. United States v. Kovel, 296 F.2d 918 (2d Cir. 1961), https://law.justia.com/cases/federal/appellate-courts/F2/296/918/131265/, requires confidential communications made to obtain legal advice with the nonlawyer necessary or highly useful to the lawyer's work. Federal Rules of Civil Procedure, https://www.uscourts.gov/sites/default/files/document/federal-rules-of-civil-procedure.pdf, supply the conditional preservation, discovery, work-product, possession/control, and sanctions framework discussed in §§4.6.1–4.6.3. Counsel-directed enterprise-vendor facts and any AI application of Kovel remain unsettled. Cited by: guide.md §§4.6–4.6.3. [primary orders, rules, and opinion]
  282. Court orders on AI and proposed Federal Rule of Evidence 707. Ropes & Gray's tracker, https://www.ropesgray.com/en/sites/Artificial-Intelligence-Court-Order-Tracker, shows varied standing orders and local rules but mixes categories, so the guide uses no count. May 2026 Evidence Rules Committee materials, https://www.uscourts.gov/sites/default/files/document/2026-05-evidence-rules-agenda-book.pdf, record the committee's May 7 decision not to advance proposed Rule 707, to revise it substantially, and to hold it for further study. The proposal is absent from the judiciary's pending-amendments list, https://www.uscourts.gov/forms-rules/pending-rules-and-forms-amendments. It has no effective date. Cited by: guide.md §4.7. [primary committee materials and current tracker]
  283. Outside counsel guidelines and generative AI. No defensible prevalence survey was located during the August 31 review. The guide therefore makes no claim about how often clients restrict AI, require consent, request vendor disclosure, or change billing rules. It retains only the operational instruction to read the actual guidelines governing the matter. Cited by: guide.md §4.8. [negative research result; no prevalence claim]
  284. Corey Epstein, ralph-methodology, https://github.com/coreyepstein/ralph-methodology, accessed August 31, 2026. Documents fresh context per task, small atomic changes, a PRD and project instructions, and automated feedback as the core Ralph mechanics. Cited by: guide.md §§1.4.2 and 5.10; companions/operator-field-notes.md. [practitioner repository; reproducible artifact]
  285. Jorge Taberóa, ralph, https://github.com/taberoajorge/ralph, accessed August 31, 2026. Concrete cross-provider loop using prd.json, prompt.md, and guardrails.md; includes stall timeout and repeated-output detection. Cited by: guide.md §§1.4.2 and 5.10; companions/operator-field-notes.md. [open-source repository; reproducible artifact]
  286. PALAN-K, llm-wiki-loop specification, https://github.com/PALAN-K/llm-wiki-loop/blob/master/SPEC.md, accessed August 31, 2026. Defines immutable raw/, model-owned wiki/, and an AGENTS.md schema; requires raw-source provenance on wiki pages and preserves outdated or disputed claims with status and history. Cited by: guide.md §§1.4.1 and 5.10; companions/operator-field-notes.md. [open-source specification; emerging implementation]
  287. r/ClaudeCode practitioner discussion, "Anyone else prompt the prompt instead of writing it?", https://www.reddit.com/r/ClaudeCode/comments/1q3p249/anyone_else_prompt_the_prompt_instead_of_writing/, captured August 31, 2026. Practitioner evidence for asking a model to draft an execution prompt before running the task. The guide treats this as a workflow signal, not proof of improved outcomes. Cited by: guide.md §§2.8 and 5.10; companions/operator-field-notes.md. [forum discussion; anecdotal]
  288. Thoughtworks, Technology Radar, Vol. 34, April 2026, https://www.thoughtworks.com/content/dam/thoughtworks/documents/radar/2026/04/tr_technology_radar_vol_34_en.pdf. Discusses the Ralph loop as an emerging agent-development technique and emphasizes short feedback cycles. Used as independent corroboration of the pattern, not of any specific implementation. Cited by: companions/operator-field-notes.md. [industry technology assessment]
  289. "LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents," arXiv:2608.18398, https://arxiv.org/abs/2608.18398, August 2026. Represents claims, actions, artifacts, and checks as typed trace records and evidence nodes for audit. Cited by: guide.md §5.10; companions/operator-field-notes.md. [research preprint; not yet peer reviewed]
  290. "PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A," arXiv:2602.21045, https://arxiv.org/abs/2602.21045, 2026. Uses structured claim-level evidence and provenance because document-level citations can hide unsupported claims and omissions. Cited by: guide.md §5.10; companions/operator-field-notes.md. [research paper; scholarly Q&A scope]
  291. Kim, E., Garg, A., Peng, K., and Garg, N., "Correlated Errors in Large Language Models," ICML 2025, arXiv:2506.07962, https://arxiv.org/abs/2506.07962, accessed August 31, 2026. Evaluates over 350 models and finds substantial error correlation: on one leaderboard dataset, models agree about 60% of the time when both err, with correlation rising as models become larger and more accurate, across providers and architectures. Frames the result as algorithmic monoculture. Cited by: guide.md §§5.6.1 and 5.10; companions/operator-field-notes.md. [peer-reviewed research]
  292. Denisov-Blanch, Y., Kazdan, J., Chudnovsky, J., Schaeffer, R., Guan, S., Adeshina, S., and Koyejo, S., "Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness," arXiv:2603.06612, https://arxiv.org/abs/2603.06612, February 20, 2026, accessed August 31, 2026. Across five models and four datasets, no tested aggregation strategy (majority vote, confidence weighting, surprisingly-popular) consistently beat a single-sample baseline despite up to 25x inference cost, and models produced correlated outputs even on random strings with no correct answer. Cited by: guide.md §5.6.1; companions/operator-field-notes.md. [research preprint]
  293. Kim, D., "Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles," arXiv:2607.20768, https://arxiv.org/abs/2607.20768, July 22, 2026, accessed August 31, 2026. In an audit of ensemble subsets of 30 models on MMLU-Pro, simple voting beat the strongest single member in only 9.98% of canonical three-model subsets. Cited by: guide.md §5.6.1. [research preprint; single author]
  294. Wittlinger, S., Meerjansen, J., Wolf, F., Wiest, I. C., Ebert, M. P., Siegel, F., and Belle, S., "Multi-LLM Disagreement as a Scalable Detector of Human Annotation Errors in Structured Data from Clinical Free-Text," medRxiv, https://www.medrxiv.org/content/10.64898/2026.05.04.26352392v1, May 6, 2026, accessed August 31, 2026. Four locally hosted open models on German-language colonoscopy reports at one institution: disagreement predicted human annotation errors at a prevalence-adjusted AUC-ROC of 0.991 (95% CI 0.987 to 0.994), and the lowest-agreement 6.5% of annotations concentrated an estimated 80% of errors. Authors state the single-center, single-language, narrow-task scope limits generalizability. Cited by: guide.md §§5.6.1 and 5.10; companions/operator-field-notes.md. [research preprint; not yet peer reviewed; narrow scope]
  295. Li, R., Gu, J.-C., Kung, P.-N., Xia, H., Liu, J., Kong, X., Sui, Z., and Peng, N., "LLM-REVal: Can We Trust LLM Reviewers Yet?," arXiv:2510.12367, https://arxiv.org/abs/2510.12367, October 14, 2025, accessed August 31, 2026. In a peer-review simulation over 100 human-authored papers and LLM-generated counterparts, LLM reviewers accepted LLM-authored papers at 78% versus 49% for human-authored papers on identical topics. Used for the proposition that a model reviewer is a biased instrument. Cited by: guide.md §5.6.1; companions/operator-field-notes.md. [research preprint]
  296. Hou, X., Wang, S., Zhao, Y., and Wang, H., "When Agents Do Not Stop: Uncovering Infinite Agentic Loops in LLM Agents," arXiv:2607.01641, https://arxiv.org/abs/2607.01641, July 2026, accessed August 31, 2026. Static analysis (IAL-Scan) evaluated on 6,549 agent code repositories across eight framework families: 74 potential findings, 68 manually confirmed unbounded-loop failures across 47 projects (91.9% precision), each a repeated path not covered by an effective bound. The paper's list of how caps fail (omitted, misused, ineffective bounds, outside the feedback path) is its motivating framing, not a measured prevalence rate. A code-analysis study, not a runtime observation study. Cited by: guide.md §5.4; companions/operator-field-notes.md. [research preprint; static analysis]
  297. "Loop Engineering: How to Design Agent Loops That Actually Converge," developersdigest.tech, https://developersdigest.tech/blog/loop-engineering-designing-agent-loops, July 12, 2026, accessed August 31, 2026. Names four convergence types (test-defined, diff-defined, count-defined, judge-defined, the last called the weakest), shows a feedback-hash stall detector escalating after two identical rounds, and requires hard ceilings on iterations, cost, and wall-clock. Cited by: guide.md §5.10; companions/operator-field-notes.md. [practitioner essay]
  298. Grigis, L., "What a ralph loop actually is, and how to run one without burning your tokens," https://lukasgrigis.dev/blog/ralph-loop/, July 1, 2026, updated August 30, 2026, accessed August 31, 2026. Argues the model that writes the code must not be the model that decides the code is done ("An agent grading its own homework gives itself an A every time"), with retry escalation to a blocked state after repeated referee failures. Cited by: guide.md §5.10; companions/operator-field-notes.md. [practitioner essay]
  299. "How OpenAI Codex implements the /goal slash command," GitHub gist, https://gist.github.com/patleeman/b1b5768393f9bf2f60865b1defeeb819, accessed August 31, 2026. Describes Codex's goal state machine: the model can create a goal and mark it complete, while pause, resume, and budget-limit transitions are system-controlled, and token budgets are enforced inline with accounting writes rather than checked separately. Hosted under a third-party account with an etraut-openai byline. Cited by: guide.md §5.10; companions/operator-field-notes.md. [informed secondary description of product internals]
  300. ChaoYue0307, awesome-loop-engineering: loop-contract schema and docs-drift example, https://github.com/ChaoYue0307/awesome-loop-engineering, accessed August 31, 2026. A JSON Schema requiring every loop contract to declare name, objective, trigger, intake, workspace, context, agents, verification, state, budget, escalation, and exit, with a dependency-free validator and a runnable docs-drift loop that logs each outcome as an append-only JSONL receipt with an evidence digest. Schema and example read end to end during this pass. Cited by: guide.md §5.10; companions/operator-field-notes.md. [open-source specification and runnable artifact]
  301. Saha, S., "What Actually Survives /compact in Claude Code," https://swapnanilsaha.com/blog/what-survives-compact-claude-code/, July 17, 2026, accessed August 31, 2026. Decay probe on Claude Haiku 4.5 via Claude Code CLI v2.1.211: ten seeded facts, source file deleted, forced compactions. Without a memory layer, zero facts survived the first boundary and 106 of 108 later summaries carried none; the two exceptions came from an accidental on-disk re-read and vanished one summarization later. With hook re-injection, ten of ten survived all 138 boundaries. The summarizer asserted it had "Maintained all 10 DECAY-FACTS" while carrying zero. Cited by: guide.md §5.5; companions/operator-field-notes.md. [single-author practitioner measurement; disclosed method and cost]
  302. Vercel (Gao, J.), "AGENTS.md outperforms skills in our agent evals," https://vercel.com/blog/agents-md-outperforms-skills-in-our-agent-evals, January 27, 2026, accessed August 31, 2026. On deliberately post-cutoff Next.js 16 APIs: 53% pass with no docs, 53% with default skills, 79% with prompt-tuned skills, 100% with an 8KB always-loaded AGENTS.md index compressed from 40KB. Vendor-run with disclosed methodology, one framework. Cited by: guide.md §5.5; companions/operator-field-notes.md. [vendor eval; methodology disclosed]
  303. He, Y., Zhao, Y., Wang, J., and Chen, H., "Is Progressive Disclosure All You Need for Long-Context Agents?," arXiv:2607.17598, https://arxiv.org/abs/2607.17598, July 20, 2026, accessed August 31, 2026. Controlled comparison of raw, flat, and hierarchical context disclosure across three harnesses and models: at twenty-book scale flat indexing raised En.QA accuracy from 0.257 to 0.462 at roughly half the cost, while hierarchical disclosure never helped and under one harness collapsed En.MC accuracy from 0.9126 to 0.6398. Effects are harness-dependent. Cited by: guide.md §5.5; companions/operator-field-notes.md. [research preprint; single lab]
  304. harrylabsj, llm-knowledge-bases-plugin, src/tools/kb_lint.ts, https://github.com/harrylabsj/llm-knowledge-bases-plugin, accessed August 31, 2026. Deterministic knowledge-base linter: recomputes each raw source's hash against the note's recorded raw_hash and flags drift, flags substantive claims paired with empty evidence sections, and flags notes older than their sources. Read end to end during this pass. Cited by: guide.md §5.10; companions/operator-field-notes.md. [open-source artifact; code verified]
  305. Glukhov, R., "LLM Wiki Maintenance: Drift, Contradictions and Review," https://www.glukhov.org/knowledge-management/knowledge-systems-architectures/compiled-knowledge/llm-wiki-maintenance-knowledge-drift/, accessed August 31, 2026. Six drift types for model-maintained wikis (source, concept, terminology, decision, citation, structure), citation drift called one of the most serious: a page cites a source but the claim no longer matches it. Proposed countermeasures include contradiction detection and risk-tiered review. Explicitly design guidance with no deployed-system measurement. The "self-citation loop" framing in the guide is the guide's own extension, not this source's. Cited by: guide.md §5.10; companions/operator-field-notes.md. [practitioner taxonomy; prescriptive, unmeasured]
  306. Chen, G., Yu, Y., and Wang, W., "Evidence-Ledger Adjudication for Claim-Evidence Traceability," arXiv:2607.26512, https://arxiv.org/abs/2607.26512, July 29, 2026, accessed August 31, 2026. An evidence-ledger agent labeling claim-evidence pairs supports, contradicts, missing, or mixed and routing unsupported claims back to authors during drafting, reaching 0.676 relation accuracy on a 2,335-row blind benchmark from AVeriTeC, CLIMATE-FEVER, and SciFact versus 0.383 for the best non-agent baseline. Independent of [289] and [290]. Cited by: guide.md §5.10; companions/operator-field-notes.md. [research preprint; measured delta]
  307. Cisco AI Threat and Security Research, "Identifying and remediating a persistent memory compromise in Claude Code," https://blogs.cisco.com/ai/identifying-and-remediating-a-persistent-memory-compromise-in-claude-code, April 1, 2026, accessed August 31, 2026. Demonstrated a malicious npm package appending attacker instructions to a Claude Code memory file during install; roughly the first 200 lines loaded into the system prompt each session, and the compromised agent persistently recommended insecure practices with no visible sign. Remediated in Claude Code v2.1.50 by removing auto-loaded user memories from the system prompt. Characterized as a supply-chain vector. Cited by: guide.md §5.8; companions/operator-field-notes.md. [vendor security research; demonstrated exploit chain]
  308. Check Point Research, "Caught in the Hook: RCE and API Token Exfiltration Through Claude Code Project Files," CVE-2025-59536 and CVE-2026-21852, https://research.checkpoint.com/2026/rce-and-api-token-exfiltration-through-claude-code-project-files-cve-2025-59536/, accessed August 31, 2026. Hook commands defined in project settings executed automatically once the general workspace-trust dialog was accepted, with no hook-specific confirmation (CVSS 8.7, fixed v1.0.111), and an ANTHROPIC_BASE_URL override redirected traffic including the Authorization header (CVSS 5.3, fixed v2.0.65). All reported issues patched before publication. Cited by: guide.md §5.8; companions/operator-field-notes.md. [primary security research; patched]
  309. Anthropic, "How we contain Claude across products," https://www.anthropic.com/engineering/how-we-contain-claude, May 25, 2026, accessed August 31, 2026. Default posture allows reads and requires approval for write, bash, and network; the OS sandbox (Seatbelt on macOS, bubblewrap on Linux) denies network by default while allowing writes inside the workspace, and that sandboxing reduced permission prompts by a reported 84%. Prompt-injection attack success on Gray Swan's Agent Red Teaming benchmark, for Claude Opus 4.7, reported around 0.1% single-attempt rising to roughly 5 to 6% after 100 adaptive attempts, given as the reason a deterministic boundary must back the model layer. Figures are vendor self-reported. Cited by: guide.md §5.8. [primary vendor engineering writeup; self-reported metrics]
  310. Khan, I., "You Don't Need Prompt Engineering Anymore: The Prompting Inversion," arXiv:2510.22251, https://arxiv.org/abs/2510.22251, October 25, 2025, accessed August 31, 2026. On GSM8K, structured "sculpting" prompting beat chain-of-thought on gpt-4o (97% versus 93%) and lost to simpler prompting on gpt-5 (94.00% versus 96.36%). Single independent author, one benchmark, one task family: suggestive, not decisive. Cited by: guide.md §§2.8 and 5.10; companions/operator-field-notes.md. [research preprint; single author, narrow scope]
  311. Hill, B., "Cost per accepted change," aifinops.dev, https://aifinops.dev/, accessed August 31, 2026. Defines the metric as proposed in The Delivery Gap (Brenn Hill, 2026) with a linked repository; the site's worked example is a generic round-figure scenario rather than reported organizational data, and no independent adoption or benchmark is claimed. Cited by: guide.md §5.10; companions/operator-field-notes.md. [single-author metric proposal; no benchmark]
  312. Vanamo, V.-M., "Token Price Is the Wrong Number: What a Merged Feature Actually Costs Across a Dozen Coding Agents," https://blog.insight-services-apac.dev/2026/07/06/cost-to-a-merged-feature, July 6, 2026, accessed August 31, 2026. One fixed feature (RFC 8628 device flow on a frozen Nuxt 4 and Drizzle repo) across 23 model-and-harness arms and 29 runs to a constant Opus 4.8 review gate, with coding and review spend tracked separately: $2.59 to $34.82 per merged result. A separate deterministic oracle caught the review gate merging code that passed as few as 12 of 38 unit tests. A "$7 to $70" range circulating in search snippets does not appear in the article. Cited by: guide.md §5.10; companions/operator-field-notes.md. [single practitioner benchmark; one task]
  313. DORA, "State of AI-assisted Software Development 2025," https://dora.dev/dora-report-2025/, accessed August 31, 2026. Reports AI adoption now correlating positively with throughput, a reversal from 2024, while delivery instability remains elevated. Direction confirmed via the report's coverage; the full PDF was not independently parsed in this pass. Cited by: companions/operator-field-notes.md. [industry research program; secondary confirmation of the specific sentence]
  314. rxdt, loopgate_harness, https://github.com/rxdt/loopgate_harness, accessed August 31, 2026, last commit August 29, 2026. Ralph-style loop harness enforcing guardrails in git hooks: forbidden files, directories, and diff patterns configured in pyproject.toml, staged changes to protected paths force-unstaged, banned patterns in added lines failed, and commits gated on a mutation-testing score. Gate, loop script, prompt, and tests read end to end during this pass. Cited by: companions/operator-field-notes.md. [open-source artifact; code verified]
  315. thu-nmrc, openloop, https://github.com/thu-nmrc/openloop, accessed August 31, 2026, last commit June 10, 2026. Loop runner with a real heartbeat file, consecutive-failure circuit breaker, OOM detection, and regex-extracted baseline gating. Negative finding: the README advertises stall detection (no output for N seconds) that the 424-line runner does not implement; the stalled field is initialized and never set. Cited by: companions/operator-field-notes.md. [open-source artifact; README claim contradicted by code]
  316. bugroo, claude-autonomy (formerly clouitreee), https://github.com/bugroo/claude-autonomy, accessed August 31, 2026, last commit March 13, 2026. Mode switcher whose "guardrails ALWAYS active" claim is backed only by a print statement; its auto profile sets Claude Code's bypassPermissions mode, which bypasses the underlying permission engine. Negative finding recorded as a due-diligence example. Cited by: companions/operator-field-notes.md. [open-source artifact; claim contradicted by code]
  317. softcane, cc-session-recover, https://github.com/softcane/cc-session-recover, accessed August 31, 2026. Hook-based session recovery that records rate-limit and quota errors with a retry window in .claude/quota-blocked.json, exposes WAIT, READY, and NONE states, and preserves task state in HANDOFF.md. Fully specified, no independent reports located. Cited by: companions/operator-field-notes.md. [open-source artifact; single tool, anecdotal]
  318. Handoff-artifact implementations: softaworks, agent-toolkit session-handoff skill, https://github.com/softaworks/agent-toolkit/tree/main/skills/session-handoff, and gsailing19, agent-handoff validator, https://github.com/gsailing19/agent-handoff, both accessed August 31, 2026. The former specifies a ten-section handoff note written for a reader who was not there. The latter's validate-handoff.py, read end to end, enforces frontmatter fields, five required sections, a literal end marker, and a .done sidecar, with typed exit codes. No production battle-testing evidence for either. Cited by: companions/operator-field-notes.md. [open-source artifacts; validator code verified, format unproven]
  319. Shipper, D., and Klaassen, K., "Compound Engineering: How Every Codes With Agents," Every, https://every.to/chain-of-thought/compound-engineering-how-every-codes-with-agents, December 11, 2025, and the EveryInc/compound-engineering-plugin repository, https://github.com/EveryInc/compound-engineering-plugin, both accessed August 31, 2026. A plan, work, assess, compound cycle in which a dedicated final command writes each cycle's lesson back into the repo as a prompt, skill, or test. Productivity claims are self-reported and unmeasured. Cited by: companions/operator-field-notes.md. [practitioner essay and plugin; self-reported outcomes]
  320. Yegge, S., "Welcome to Gas Town," https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16dd04, January 1, 2026; Appleton, M., independent analysis, https://maggieappleton.com/gastown; and Hacker News discussions 46458936 and 46734302, all accessed August 31, 2026. Role-hierarchy orchestration (orchestrator, disposable workers, supervisor, merge-queue owner) over a git-backed task store. Documented problems: author-admitted complexity, overwhelming onboarding, reported costs in the thousands per month, and reviewer capacity as the unsolved bottleneck. Cited by: companions/operator-field-notes.md. [practitioner primary account with independent critique; mixed field reports]
  321. Wu, W., "When Errors Become Narratives: A Longitudinal Taxonomy of Silent Failures in a Production LLM Agent Runtime," arXiv:2606.14589, https://arxiv.org/abs/2606.14589, June 15, 2026, accessed August 31, 2026. Single-system eight-week study of 22 postmortemed incidents: replayed failure-derived guardrails showed 0% ex-ante prevention and 87% regression blocking, with roughly 70% of silent failures first detected by a human. Per-category tables could not be extracted during this pass. Cited by: companions/operator-field-notes.md. [single-system preprint; aggregate figures only]
  322. Microsoft, "Shadow Mode in Case Management Agent: Validate Before You Automate," Dynamics 365 blog, https://www.microsoft.com/en-us/dynamics-365/blog/it-professional/2026/07/09/shadow-mode-case-management-agent/, July 9, 2026, accessed August 31, 2026. The agent predicts, recommends, and drafts against live cases while sending nothing and changing no records. No published promotion thresholds or measured results. Cited by: companions/operator-field-notes.md. [vendor product announcement; existence proof only]
  323. Prompt-change CI gating documentation: Promptfoo, "CI/CD Integration," https://www.promptfoo.dev/docs/integrations/ci-cd, and Langfuse, "Prompt CI/CD" and "LLM regression testing," https://langfuse.com/resources/engineering/prompt-cicd and https://langfuse.com/resources/engineering/llm-regression-testing, all accessed August 31, 2026. Two independent tool ecosystems converge on golden-dataset evals with pass-rate thresholds failing CI, immutable prompt versions, and weighted rollout. Vendor documentation; no independent defect-catch-rate measurement located. Cited by: companions/operator-field-notes.md. [vendor documentation; convergent implementation shape]
  324. Kenny, J., "Solving Agent Context Loss: A Beads and Claude Code Workflow for Large Features," https://jx0.ca/solving-agent-context-loss/, January 2, 2026, accessed August 31, 2026. Git-backed dependency-linked task graph with one fresh subagent per task and two-stage review. One independent critical account reports the agent mishandling merge conflicts and tool bugs confusing the agent. Cited by: companions/operator-field-notes.md. [practitioner writeup; adoption signals with an unresolved reliability complaint]
  325. Imbad0202, academic-research-skills, v3.21.1, https://github.com/Imbad0202/academic-research-skills, August 24, 2026, accessed August 31, 2026. Cross-model verification with blind disagreement checkpoints that withhold output pending explicit human decision, and an author warning that two models seeing the same prompt structure is not independence. Unmeasured for error-catching efficacy. Cited by: companions/operator-field-notes.md. [open-source artifact; shipped human-halt design]
  326. Anthropic, "How we built our multi-agent research system," https://www.anthropic.com/engineering/multi-agent-research-system, June 13, 2025, accessed August 31, 2026. Agents typically use about 4x more tokens than chat and multi-agent systems about 15x; the multi-agent system beat single-agent Opus 4 by 90.2% on Anthropic's internal research eval; Anthropic cautions the shape fits parallelizable research and not most coding tasks. Cited by: companions/operator-field-notes.md. [primary vendor engineering writeup; internal eval]
  327. AI Incident Database, Incident 1193 (Deloitte Australia welfare-compliance report), https://incidentdatabase.ai/cite/1193/, accessed August 31, 2026. Deloitte issued a partial refund to the Australian government on a roughly A$440,000 report after fabricated references and a misattributed Federal Court judgment were found and the firm acknowledged generative AI use in its preparation. Cited by: companions/operator-field-notes.md. [incident record aggregating primary reporting]
  328. STAT News, "Lancet study finds steep rise in fraudulent citations in academic papers," https://www.statnews.com/2026/05/07/lancet-study-finds-steep-rise-fraudulent-citations-academic-papers/, May 7, 2026, accessed August 31, 2026. Columbia researchers analyzing over two million papers and 97 million citations found the share of papers with fabricated references rose from 1 in 2,828 in 2023 to 1 in 458 in 2025 and 1 in 277 in the first seven weeks of 2026. Cited by: companions/operator-field-notes.md. [secondary reporting on a Lancet-published analysis]
  329. atomicstrata, llm-wiki-compiler, https://github.com/atomicstrata/llm-wiki-compiler, accessed August 31, 2026. Independent knowledge-lint implementation (llmwiki lint): broken citations validated against source line ranges, freshness and orphan flags, contradiction detection, schema compliance, usable as a CI gate. Author-documented limits: weak fit when the domain changes faster than review can validate. Cited by: companions/operator-field-notes.md. [open-source artifact; self-documented scope limits]
  330. Crosley, B., "Context Compaction Is a Decision, Not a Threshold," https://blakecrosley.com/blog/agent-context-compaction, June 23, 2026, accessed August 31, 2026. Argues for rubric-gated compaction exposed as a callable action. Relays results from a paper this pass could not locate independently, so its numbers are not citable. Cited by: companions/operator-field-notes.md. [practitioner essay; underlying paper unlocated]
  331. Pan, T., "Consensus Protocols for Multi-Agent Decisions: What Happens When Your Agents Disagree," https://tianpan.co/blog/2026-04-12-consensus-protocols-multi-agent-decisions-when-agents-disagree, April 12, 2026, accessed August 31, 2026. Synthesis recommending a two-to-three round debate cap before forcing a decision. Its quantified claims could not be traced to underlying studies during this pass and are treated as unverified. Cited by: companions/operator-field-notes.md. [practitioner synthesis; numbers unverified]
  332. Prediction Guard, "Least agency: a blast-radius governance framework for persistent AI agents," https://predictionguard.com/blog/least-agency-blast-radius-governance-framework-persistent-ai-agents, updated August 28, 2026, accessed August 31, 2026. Scores agent blast radius by identity reach, data-access scope, and downstream connectivity, with tiered approvals. Self-flagged as providing no quantified caps. Cited by: companions/operator-field-notes.md. [consultancy framework; unmeasured]