Layman's Terms← All companies

scale.com

Scale AI

Last reviewed June 2026 — a point-in-time snapshot.

TL;DR

1 / 3

What Scale AI actually does

Scale AI builds the training data that frontier AI models learn from, the largest general-purpose models behind tools like ChatGPT. It also runs the tests those models are graded on, and sends teams into hospitals, banks, and militaries to build AI that has to work. The company says 90% of leading model builders run on its data, and one in four of its data contributors holds an advanced degree.

2 / 3

The analogy

Michelin made tires. In 1900 it started rating restaurants to sell more of them, and a century later its stars are the standard every ambitious kitchen chases, awarded by a company that doesn't cook. Scale's evaluation lab does this for AI. Its tests, like Humanity's Last Exam, are where labs go to judge how good a new model is. Scale supplies the data behind the models and set the scoreboard they're ranked on.

3 / 3

What only Scale can claim

Scale grades every frontier model, then puts its name on building the ones sent to work where a wrong answer is expensive.

  • It published 450+ model evaluations in 2025. The payoff: when Scale names a model for your task, it has already measured every serious rival.
  • Its teams embed on-site with customers like Mayo Clinic and BP and own the result, not the software. Why a buyer cares: most company AI projects stall, and Scale ships the ones that don't.
  • It runs AI for the Pentagon and for governments in Qatar and the UAE. What this means: Scale works in secured rooms most vendors can't enter.

The Full Read

Meta bought 49% of Scale AI, and within days two of its biggest customers started walking. Scale still booked its biggest year ever.

What Scale AI actually does

Scale AI does three things, and they feed each other. First, it produces the data that frontier AI models train on, the largest general-purpose models behind tools like ChatGPT. That means paying experts to write, rate, and correct millions of examples so a model learns to reason instead of guess. Scale says 90% of the world's leading model builders run on its data, and a quarter of its contributors hold advanced degrees. Second, through a lab it calls SEAL, it builds the tests those models are graded on, including one named Humanity's Last Exam. Third, and growing fastest, it sends teams into companies and governments to build and run AI in the real world. These teams pick the use case, build the system, and stay on the hook for whether it works. Hospitals, energy companies, and defense agencies are the customers. The older way was buying raw data labeling by the hour and hoping the model held up once it shipped.

The analogy

Michelin made tires. In 1900 it began rating restaurants to get people to drive more, and over a century the stars became the standard every ambitious kitchen organizes around, awarded by a company that doesn't cook. Scale's SEAL lab plays that role for AI. Scale doesn't build the frontier models, but its tests are where labs and developers go to see how a new release actually performs. The mechanism is the same. An outsider's scorecard becomes the thing the whole field is measured against, and the scorekeeper earns an authority no single competitor can match.

Who they serve

Scale serves three buyers, and the mix shifted hard in 2025. AI labs and model builders buy its data and its evaluations. Companies buy its applications work: Mayo Clinic, BP, Allianz, Cengage, Howard Hughes, and Time are named on its site. Governments are the fastest-growing slice, from the US Department of Defense to sovereign programs in Qatar, the UAE, and Saudi Arabia, where a government runs AI it controls on its own soil. The split tells you who Scale is really built for now. After Meta took a large stake in 2025, several frontier labs that compete with Meta pulled back, and Scale's growth tilted toward companies and governments, buyers who care less about who owns a piece of Scale than about whether the system works.

Who shouldn't use it

Scale isn't the right call for everyone. A frontier lab that competes with Meta has reason to keep its plans away from a company Meta part-owns, and rivals like Surge AI or Mercor do the same data work without that complication. A team that just needs cheap, high-volume basic labeling will pay less with Sama, Appen, or a self-serve tool like Labelbox. A company with its own mature data team and annotator network is paying for managed expertise it already has. And a small startup on a tight budget is not who Scale is built around; its contracts are sized for companies and governments, not low-commitment self-serve.

Sample customer stories

Hypothetical example: a regional health system. It has tried three AI vendors to summarize patient charts. Each demo dazzled, and each pilot died when doctors found the summaries missed the one detail that mattered. Scale's applications team doesn't start with a model. It sits with the clinicians, narrows to the charts where summarizing is safe, builds the system, and stays accountable for accuracy once it's live. What no longer happens: a tool gets bought, goes unused, and quietly gets written off.

Hypothetical example: a frontier lab. It needs ten thousand graduate-level physics problems with verified solutions to push its model's reasoning, and it can't hire that expertise fast. Scale already has the contributors, a quarter of them with advanced degrees, and the workflow to produce the data at the bar the lab demands, then grade the result on SEAL.

What only Scale can claim

Scale is the only company that grades every frontier model and then stakes its own name on building the ones that go to work in the highest-stakes places.

The grading is real and public. In 2025 its SEAL lab shipped 15 new tests and more than 450 evaluations across 50-plus model releases, and benchmarks like Humanity's Last Exam are now cited across the industry. The payoff for a buyer: when Scale tells you which model to trust for a task, it has already measured every serious alternative against it.

The building is hands-on. Scale's applications teams move in on-site, pick the use case that can actually work, and own the outcome instead of shipping software and leaving. Why this matters: most company AI projects stall before they reach production, and Scale is paid to deliver the ones that don't.

The flywheel connects the two. Grading every model tells Scale, task by task, what is reliable. Deploying in production tells it what breaks in the real world. Each makes the next test and the next deployment sharper, and no competitor sits on both halves of that loop.

The hardest rooms are a moat of their own. Scale runs AI for the Pentagon, alongside Anduril and Microsoft, and full sovereign AI programs in Qatar, the UAE, and Saudi Arabia. What this means: security clearances, classified settings, and government trust are slow to earn and hard for a rival to copy.

Data, evaluations, and deployment are sold separately, so a customer can start with one and pull in the others as the work gets harder. The option to go deeper is the point.

Why this is hard, and why it matters now

Durable structural shift: AI has moved from demos to decisions that carry real consequences, a misread scan, a bad targeting call, a wrong loan. Once a model shapes those calls, the binding question stops being whether you can get a model and becomes which one you can trust, and who is accountable when it's wrong. That is a testing-and-accountability problem, and it's the one Scale is built around.

Current shift: through 2025 and into 2026, boards started demanding returns on their AI spending, and most pilots weren't delivering. Governments moved past experiments and began buying AI they own and run outright. At the same time the market for training data fractured: after Meta took its stake, labs that compete with Meta grew wary of sharing plans with Scale. That pushed Scale toward companies and governments, buyers for whom Meta's ownership simply isn't the deciding question.

What people would use instead

Without Scale, the work splits across familiar fallbacks. For data, labs increasingly route to Surge AI and Mercor, the same expert work without the Meta tie. For basic labeling, teams use Labelbox or their own annotators. For deployment, companies hire Palantir, whose engineers also embed and own outcomes, or lean on Accenture and Deloitte, who already sit inside their accounts. For deciding which model to trust, some skip SEAL and use open tests or crowd-ranked arenas. The common fallback across all three is building in-house: labs and large companies with their own machine-learning teams often do the work themselves. The catch is that few can do all three at once, which is the seam Scale fills.

Competition

The field splits in two.

  • Data and evaluations: Surge AI, Mercor, Turing, Micro1, Labelbox, Sama, Appen, and Snorkel AI; on the testing side, Artificial Analysis and crowd-ranked arenas.
  • Applications and government: Palantir, the consultancies Accenture, Deloitte, and Booz Allen, and the cloud providers' own services from Microsoft, Google, and Amazon.

Surge AI is built around the same expert data Scale sells and, though bootstrapped, reportedly passed Scale on revenue in 2024. Mercor is built around a marketplace of vetted experts and scaled fast in 2025. Palantir is built around forward-deployed engineers who own deployment in companies and defense, the closest match to Scale's applications business. Scale is built around holding both halves at once, the data and tests that judge models and the teams that ship them, which is the seam none of the others sit across.

For the Team

Website analysis

The finding is a buried moat. The homepage gives roughly equal space to two things, the tests Scale runs on frontier models and the teams it sends into companies, and it never connects them. The connection is the whole point. Grading every model tells Scale what's reliable for which task; deploying in production tells it what breaks; each sharpens the other. No competitor holds both halves. That loop is Scale's most defensible asset in a year when its old moat, deep ties to the frontier labs, was the part the Meta deal eroded. The homepage also leads with an abstract line, "The world's most important decisions need reliable AI systems," which states a need rather than what Scale does. For a company this well known, that's a missed chance to claim the position only it can hold, and to absorb the elephant in the room: the neutrality question that pushed Google and OpenAI away.

Website rewrite

This is a hero-broken, subhead-fixable case. The hero states a need instead of a position; the subhead's opening line is strong brand voice worth keeping, and the rest leans on a vague "AI stack" claim.

  • Current hero (verbatim): "The world's most important decisions need reliable AI systems."
  • Current subhead (verbatim): "Reliable AI has no shortcuts. Scale works across the AI stack, from the data that trains the models you rely on, to the systems that put them to work. Humans stay in the loop."
  • Current CTA (verbatim): "Book demo"
  • Rewritten hero: We build the AI behind decisions that can't be wrong.
  • Rewritten subhead: Reliable AI has no shortcuts. A quarter of our data contributors hold advanced degrees, and our tests are where labs judge new models. Then our teams build and run what you put into production, people in the loop.
  • Rewritten CTA: Keep as is.
  • Reasoning: the hero now claims what Scale does while keeping the "important decisions" equity it has earned; the subhead trades the vague "AI stack" line for two concrete proofs (expert contributors, the tests) plus the embedded-team pillar; the CTA already converts and doesn't echo the subhead's action. Length parity holds: hero 9 words to 10, subhead 33 to 38. Job-to-be-done is positioning, not awareness, because the homepage's visitor already knows Scale.

Messaging to consider

  1. When the answer has to be right, you don't buy a model. You hire the people who grade them all.
  2. The data behind 90% of frontier models. The teams behind the missions that run them.
  3. Most company AI never ships. We're paid for the part that does.

Lead with the first. It passes ownability most cleanly: only Scale grades every frontier model, so only Scale can credibly reframe the decision from "pick a model" to "hire the graders," and it moves Scale out of the crowded "AI vendor" category.

Likely next questions a prospect would have

  • After the Meta deal, how is my data kept away from Meta, and what is the guarantee in writing?
  • What does "own the outcome" mean in a contract: fixed fee, service levels, and who owns the resulting software?
  • Is Scale enterprise-only now, or is there a smaller-scale option, and what are the minimums?
  • For government and regulated work, which security certifications and data-residency options apply?
  • How independent is SEAL, given Scale sells data to the same labs it tests and Meta owns 49%?
  • What is a realistic timeline from first conversation to a working, deployed application?

Sources

Company sources: https://scale.com/ , https://scale.com/blog/scales-next-era-building-for-2026 , https://scale.com/blog/scale-ai-announces-next-phase-of-company-evolution , https://labs.scale.com/leaderboard , https://scale.com/leaderboard/humanitys_last_exam

Third-party sources: https://www.cnbc.com/2025/06/12/scale-ai-founder-wang-announces-exit-for-meta-part-of-14-billion-deal.html , https://fortune.com/2025/06/13/meta-scale-ai-alexandr-wang-superintelligence-team/ , https://www.cnbc.com/2025/06/14/google-scale-ais-largest-customer-plans-split-after-meta-deal.html , https://techcrunch.com/2025/06/18/openai-drops-scale-ai-as-a-data-provider-following-meta-deal/ , https://techcrunch.com/2025/08/29/cracks-are-forming-in-metas-partnership-with-scale-ai/ , https://www.inc.com/jennifer-conrad/surge-ai-edwin-chen-scale-ai-meta-alexandr-wang/91204563 , https://sacra.com/c/surge-ai/ , https://www.axios.com/2026/05/06/scale-ai-jason-droege-reliable-ai , https://en.wikipedia.org/wiki/Scale_AI