← Read

AI · · 7 min read

The browser has a language model, and almost no promises about it

The Prompt API puts a LanguageModel class in the page, with structured output, tool calls and a context window you manage. What it guarantees is the interesting part.

One class, no key, no endpoint:

const session = await LanguageModel.create();
const result = await session.prompt("Write me a poem.");

The Prompt API is the general-purpose member of the browser's built-in AI family — the one where you do your own prompt engineering instead of calling a task-specific API. It is being developed by the Web Machine Learning Community Group, and at the time of writing MDN's compatibility data records LanguageModel in Chrome 148 on desktop, in Edge 138 behind the #edge-llm-prompt-api-for-phi-mini preference, and nowhere else — not Firefox, not Safari, not Chrome on Android.

The availability dance and the download monitor work the same way as in the Translator API, so this article covers what is different: what the explainer promises, structured output, tool calls, and the context window you are expected to manage yourself.

github.com ↗Explainer for the Prompt APIThe proposal repository, with the full API surface, the non-goals and the open issues. Source: W3C Web Machine Learning Community Group

Read the non-goals before you design around it

Most APIs bury their limits. This one states them, and they should shape what you build:

  • "We do not intend to force every browser to ship or expose a language model… It would be conforming to implement this API by always signaling that no language model is available."
  • "it may also be viable to implement this API entirely by using cloud services instead of on-device models."
  • "We do not intend to provide guarantees of language model quality, stability, or interoperability between browsers."

So: the model may not exist, it may not be on the device, and two browsers that both implement the API may give you materially different answers to the same prompt. Knowing or controlling whether a session runs on-device is listed as a potential goal the group is "not yet certain of", alongside exposing a model identifier.

The practical reading is that this is a progressive enhancement for features that degrade gracefully — suggesting tags, drafting a first pass, classifying text you can also classify another way. It is not a substitute for a server model where the output has to be consistent, and it is not yet a privacy guarantee you can make to users in writing.

One thing it does intend: the models are instruction-tuned, not base models. A prompt like "Write a poem about trees" should produce a poem, not more instructions.

Structured output is the reason to reach for it

A model that returns prose is awkward to program against. responseConstraint takes either a JSON schema object or a RegExp:

const schema = {
  type: "object",
  required: ["rating"],
  additionalProperties: false,
  properties: {
    rating: { type: "number", minimum: 0, maximum: 5 }
  }
};

const result = await session.prompt(
  "Summarize this feedback into a rating between 0-5: " +
  "The food was delicious, service was excellent, will recommend.",
  { responseConstraint: schema }
);

const { rating } = JSON.parse(result);

The errors are specific, and worth handling separately:

SituationWhat you get
Valid schema using features this browser doesn't supportNotSupportedError DOMException
The model can't produce output matching the schema or RegExpSyntaxError DOMException
The value is neither a RegExp nor a valid JSON schemaTypeError

The part that catches people: by default the schema is sent to the model as part of the message, so it consumes context window. You can measure that cost by passing responseConstraint to session.measureContextUsage(). If the schema is large and you would rather not pay for it, omitResponseConstraintInput: true suppresses it — but then the explainer strongly recommends describing the shape you want in the prompt text, because the model no longer sees it. Passing that option without responseConstraint is a TypeError.

There is a lighter-weight nudge too. Adding prefix: true to a trailing "assistant"-role message prefills the start of the response, so the model continues from it rather than warming up:

const sheet = await session.prompt([
  { role: "user", content: "Create a JSON character sheet for a gnome barbarian" },
  { role: "assistant", content: '{\n  "name": ', prefix: true }
]);

prefix anywhere other than a final assistant message throws a SyntaxError DOMException.

Tool calls run in your page, sometimes several at once

The tools option lets the model call your JavaScript. Each tool is a name, a description, an inputSchema, and an async execute() that the browser invokes and whose return value it feeds back to the model:

const session = await LanguageModel.create({
  expectedInputs:  [{ type: "text", languages: ["en"] }, { type: "tool-response" }],
  expectedOutputs: [{ type: "text", languages: ["en"] }, { type: "tool-call" }],
  tools: [{
    name: "getWeather",
    description: "Get the weather in a location.",
    inputSchema: {
      type: "object",
      properties: { location: { type: "string", description: "The city to check." } },
      required: ["location"]
    },
    async execute({ location }) {
      const res = await fetch("https://weatherapi.example/?location=" + location);
      return JSON.stringify(await res.json());
    }
  }]
});

const result = await session.prompt("What is the weather in Seattle?");

The explainer flags the behaviour that will break a naive implementation: your execute() may be called multiple times concurrently. Ask which of Seattle, Tokyo and Berlin is warmest and getWeather runs three times in parallel, with the model waiting on "the equivalent of Promise.all()" before composing its answer. Anything your tool does that is not safe to run concurrently — a shared mutable cache, a rate-limited endpoint, a write — needs handling inside execute().

The context window is yours to manage

Sessions accumulate. Two read-only properties tell you where you are:

console.log(`${session.contextUsage} used of ${session.contextWindow}`);

measureContextUsage() costs you a measurement rather than a turn, accepts the same inputs as prompt() including multimodal arrays, and can be aborted with an AbortSignal. The actual tokenisation is deliberately not exposed, so you cannot count tokens yourself — and implementations must include control tokens in the count, so your own estimate would be wrong anyway.

When a prompt does not fit, the session does not fail first. It evicts: the oldest prompt/response pairs are dropped one at a time until the new prompt fits, and initialPrompts are never removed. That is a quiet correctness bug waiting to happen, so listen for it:

session.addEventListener("contextoverflow", () => {
  // earlier turns have just been dropped
});

If eviction cannot free enough room, prompt() rejects with a QuotaExceededError carrying requested and the available context window, and nothing is removed. append() — which sends messages ahead of time so the model can start processing them, useful for images — can trigger the same eviction and the same event.

If you have followed this API for a while, note the renames: inputUsage became contextUsage, inputQuota became contextWindow, measureInputUsage() became measureContextUsage(), and onquotaoverflow became oncontextoverflow. The old names are deprecated.

Sampling modes replaced temperature

temperature and topK are deprecated for web pages. Passing them logs a deprecation warning and is ignored at runtime, and the corresponding session properties come back undefined. LanguageModel.params() is extension-only now.

The replacement is a named scale: "most-predictable", "predictable", "slightly-predictable", "balanced" (the default), "slightly-creative", "creative" and "most-creative". The resolved value is readable back off the session as samplingMode. In the contexts where raw parameters still work, passing both a samplingMode and a temperature or topK rejects create() with a TypeError.

Where it runs

By default the API is exposed to top-level windows and their same-origin iframes. A cross-origin iframe needs delegation:

<iframe src="https://example.com/" allow="language-model"></iframe>

It is not available in workers at all, which the explainer attributes to the difficulty of establishing a responsible document for the permissions-policy check. So the long-running generation you wanted to push off the main thread stays on the main thread. Use promptStreaming() and render as chunks arrive.

FAQ

Is this the same as Chrome's built-in AI?

The Prompt API is the general-purpose one. The task-specific APIs — translator and language detector, summarizer, writer and rewriter, proofreader — are separate proposals from the same group. Prefer those where they fit: the explainer's own framing is that the Prompt API buys you more capability "at the cost of requiring them to do their own prompt engineering".

Does the model run on the device?

Not necessarily. The explainer says implementing the API entirely with cloud services would be conforming, and there is currently no way to ask for on-device only. Do not promise users that their text stays local.

How do I know a session will work before I create one?

LanguageModel.availability() resolves to "unavailable", "downloadable", "downloading" or "available" for the options you pass it. Anything other than "available" means you should tell the user something is about to be fetched.

Can I reuse a warmed-up session?

Yes — sessions persist their interactions, and can be cloned so you branch from a common set of initial prompts rather than re-processing them. destroy() releases one.

Is it standardised?

Not yet. It is a Community Group proposal in active development, with implementations described in the explainer as "experimentally available" in Chrome and Edge. Expect the surface to keep moving; it already has.

Sources

More to read