SproutVest Field Notes · No. 2 · August 2026
Which AI model should your business actually run?
The first two sections are free and below, in full. The complete guide, with the cost model and the August pricing landscape, is $199.
Or jump to the full guide, $199 →
Most model selection begins with a leaderboard and ends with a bill nobody forecast. The benchmark used to make the decision measured something your business does not do, and the pricing was compared per million tokens when your actual unit is a resolved customer request. Meanwhile, switching costs were assumed to be near zero right up until the day came to make the switch.
This guide is about the business decision, not the scorecard used to measure it. It assumes you are running a business somewhere between fifty and a thousand people, that you have one to five places where a model would plausibly earn its keep, and that you do not have a research team to run a six-month AI bake-off. It assumes you will have to defend your choice to a CFO who does not care which model won a math competition.
Sections 1 and 2 are below in full. No email required, nothing to unlock.
1. Why model choice is actually the second question
The first question you should ask is what is actually the work, or 'job to be done' (JTBD). Model selection follows from a clear description of the task, and almost every expensive mistake in this area traces back to picking the model before describing what the job was.
Three patterns produce most of these mistakes.
Picking on a benchmark that does not resemble your work. Public benchmarks measure hard, general, adversarially selected problems, because that is what makes them useful for comparing frontier research. Your support triage workload is none of those things. A model that scores three points higher on a graduate-level reasoning benchmark may perform identically on your task, or worse, if your task rewards instruction-following and format discipline over reasoning depth. The leaderboard is measuring a different sport.
Standardizing one model across four different jobs. Organizations reach for a single vendor because procurement is easier and the architecture diagram looks tidier. Then the same model that is comfortably overqualified for classifying inbound emails is underqualified for drafting a technical response, and you are simultaneously overpaying on one task and underperforming on another. Standardization is a real benefit and it is not free.
Treating the decision as permanent. This market reprices and re-releases on a cadence measured in months. A choice that assumes a three-year horizon will be wrong long before the horizon arrives. The correct posture is a default model, an escape hatch, and a low enough switching cost that revisiting the decision is a Tuesday rather than a project.
That last point is the reframe worth carrying through the rest of this guide. You are not selecting a model. You are selecting a default, an escalation path for the cases the default handles badly, and an architecture that lets you replace either one without a rewrite.
2. The seven criteria that actually move cost and risk
Everything that matters about a model, for a business decision, reduces to seven questions. Each one has a test you can run without a data science team.
Task fit
Does the model do your specific job well, measured on your data.
The test: assemble one hundred real examples from last month's actual work, with known-good outcomes attached. If you cannot build that set in a day, you do not yet understand the task well enough to choose a model for it, and no amount of vendor comparison will rescue you. This set is the single most valuable artifact you will build in this process, and Section 7 is about using it.
Latency tolerance
How long the answer can take before the value degrades.
The test: ask what the human does while waiting. If they sit and watch a cursor, you are buying 95th percentile response time and you should be measuring it under concurrency, not in a quiet demo. If the job runs overnight in a batch, latency is nearly free and you are shopping on price alone. These two cases lead to different models and different prices, and teams routinely buy the expensive answer for the overnight job.
Context requirements
How much input the model must hold at once.
The test: measure the actual 95th percentile input size in your real traffic. Most teams buy for an imagined worst case that occurs in under one percent of requests, and pay for that headroom on every single call. If the tail genuinely needs a larger window, route the tail to a different model instead of buying the ceiling for everything.
Tool and integration needs
Whether the model has to call your systems, and how reliably.
The test: count how many of your tasks require the model to invoke something rather than just produce text, and whether it needs to chain several calls or make exactly one. Single-call tool use is close to a solved problem. Multi-step chains where each step depends on the last are where models diverge sharply, and where a cheaper model quietly costs more in failed runs.
Data residency and governance
What you are permitted to send, and where it may be processed.
The test: name the strictest customer contract you hold and read what it says about subprocessors and data location. That contract, not your own comfort level, sets the boundary. If you cannot add a new subprocessor without customer notification, that is a procurement timeline, and it belongs in the plan before you fall in love with a vendor.
Switching cost
How long it would take to replace the model behind your product.
The test: ask your engineering lead how many days it would take to swap the model and re-validate. If nobody knows, the honest answer is longer than you think, and reducing that number is worth more than a marginal capability gain. Section 8 covers how to keep it low.
Total cost per resolved task
What one unit of finished work actually costs.
This is the criterion most often measured wrongly, and it gets its own section.
What the other six sections tell you
You now have the criteria. What you do not have is any of the arithmetic, and the arithmetic is where the counterintuitive results live. Three of them, stated plainly so you can judge whether the rest is worth $199:
Cheap models are usually the expensive choice. Per-attempt inference costs across the current market span roughly 56x. Put one support agent next to that spread and it collapses into rounding error. Section 4 shows that upgrading from the cheapest open model to the most expensive frontier model pays for itself if the frontier model resolves one third of one percentage point more tickets without a human. Not three points. One third of one point.
Identical open weights cost up to 10x more depending on who serves them. The same Llama 3.3 70B weights are priced $0.10 per million input tokens at one host and $1.04 at another, on the same day. Section 6 has the full spread across providers, which means the first cost review of any existing deployment should check the host before it checks the model.
Self-hosting almost certainly costs you more. Section 3 works the break-even between renting a GPU and paying per token. It lands near 9.5 million requests per month, sustained around the clock. Below that, and nearly every mid-market workload is far below it, renting hardware is the more expensive option and the gap is not close.
Inside the paid edition
3 · Open or closed, and what each really commits you to
Including the utilization break-even worked with current GPU and per-token rates, and the middle path most teams should take that satisfies neither side of the usual argument.
4 · Cost per resolved task
The only unit that matters, with the full cost table, the break-even deflection math, and the three conditions that invert the conclusion.
5 · The decision framework
Mapping workload against the constraint that actually dominates, and why the right answer is usually two models rather than one.
6 · The model landscape, August 2026
Every commonly available closed and open model with input and output pricing, context windows, and deployment notes. Anthropic, OpenAI, Google, DeepSeek, Qwen, Llama, Gemma, and the GPT-OSS line, plus hosted-provider price spreads, self-hosting GPU rates, and the gateway category. Every figure cited to the vendor's own page with its retrieval date.
7 · How to evaluate in two weeks
Without a data science team, including the sample size that makes a result meaningful and the failure modes that only appear under production load.
8 · What to revisit in six months
The contract terms worth renegotiating and the three habits that keep switching cheap.
$199. A dated PDF edition, delivered by email. Edition 1, August 2026.
If it does not change a decision you were about to make, reply to your receipt within 7 days and I will refund it in full, no questions.
Delivery is by email from [email protected], usually within a few hours and always within one business day. If your copy has not arrived, write to [email protected] and it will be sent straight away.
Section 6 is a dated snapshot and will be out of date within months. It says so in the guide. Sections 3 through 5, 7, and 8 are the durable part and the reason to buy this.
Not ready to buy?
Paid editions are published as the market moves enough to justify a new edition. Leave your email and you will get the next one's release note, plus every free edition as it ships.
You are on the list. You will get the next edition's release note at that address. Questions in the meantime: [email protected].
Disclosure
Erick Watson was Co-Founder and CFO of Chainlodge, a data center campus targeting AI compute workloads, which wound down in 2026. Sections 3 and 4 of the paid guide discuss self-hosting and compute sourcing, which is a market SproutVest has a commercial interest in. You should know that before you read the recommendations, not after.
Who wrote it
Erick Watson has shipped AI products into production where being wrong was expensive. He led product for an AI valuation platform that ran inference against all 105 million US households and holds #1 AVM accuracy, built data products serving 34M+ members at Elevance Health, and through Chainlodge saw today's AI compute market from the supply side. This guide is the selection conversation he has had with operators repeatedly, written down once.
