Vercel has published a detailed guide on integrating GPT-6 Sol into application workflows through AI SDK's experimental evaluation API, a feature that lets developers ask named questions about shared data at specific points in an application's execution.

What the Evaluation API Does

The AI SDK evaluation API allows application code to pose structured questions about a piece of shared input called the state. In a support application, for instance, the API can assess a draft reply against the customer's original message and the relevant policy before that reply reaches a human reviewer. The evaluation happens during the application's workflow, not after deployment.

The OpenAI evaluation adapter behind this feature uses the Responses API with structured output to produce three types of answers: Choice, Score, and Boolean probability. Choice answers select from predefined options like ready, revise, or needs_review. Score answers return a numeric value against a rubric. Boolean probabilities estimate whether a specific statement is true. Notably, Choice and Score answers do not include probability distributions, and Boolean results are generated estimates — a score of 0.9 should not be interpreted as proof that the model is correct nine times out of ten.

A Practical Example: Reviewing a Customer Reply

The guide walks through a concrete scenario. A customer asks whether an unused purchase qualifies for a return. A draft reply claims the refund has already been processed. Even if the customer is eligible, the policy alone cannot support a factual assertion that money has been returned. The evaluation separates whether the reply addresses the request from whether its specific claims are supported.

Developers supply the customer message, the applicable policy, and the draft reply as the state. If a payment record exists, it is included as a distinct source. Keeping inputs separate defines what the model should judge and gives reviewers a way to investigate disputed results.

The setup requires three steps. First, install the AI SDK and OpenAI provider with npm install ai @ai-sdk/openai and configure the OPENAI_API_KEY environment variable. Second, define the evidence and questions in a shared module. Third, call the experimental_evaluate function with the model, state, questions, and optional reasoning effort settings.

Choosing Between GPT-6 Models

The evaluation adapter works with three GPT-6 variants, each suited to different workloads.

Sol handles the general evaluation case and supports reasoning effort from none through max. The default is no reasoning, but the guide's example selects medium for a balanced review. Luna fits focused evaluations repeated across many inputs — a narrow rubric applied to high volume. Luna also supports none as its minimum effort. Astra is the candidate for complex reasoning across several records, but it does not support reasoning effort set to none. If no effort is specified, the OpenAI provider strips the setting and returns an unsupported-setting warning, so developers must explicitly set a supported level like high.

For a fair comparison across models, the guide recommends using the same supported effort setting on all three, then tuning each model separately based on results.

Validating That Evaluations Are Actually Useful

The API is experimental and can change in patch releases. The guide emphasizes that measuring how accurately Sol evaluates requires a separate comparison against examples that human reviewers have already judged.

Teams should build a test set of draft replies with expected dispositions, including cases where a reply is fluent but unsupported and cases where the policy itself is incomplete. False approvals and unnecessary revision requests should be tracked separately because they have different consequences. After changing the model, reasoning effort, or question definitions, compare the model's judgments against those labels. Record request duration and token usage alongside the decisions to assess cost at expected volume.

The evaluation returns one complete result and does not stream answers. An array passed as state serves as shared context for a single evaluation, such as a conversation history. Unrelated replies should be evaluated separately so each decision has an unambiguous subject.

Handling Failures and Limitations

The OpenAI adapter fails the evaluation call entirely when it receives a refusal, truncated output, or invalid answers. That failure is distinct from a completed evaluation whose disposition is needs_review. The former means no usable result exists; the latter is a judgment about the evidence. Applications should record failures separately, use abortSignal for deadlines or cancellation, and configure maxRetries for transient provider errors. Retries recover from temporary errors, but missing evidence must be added to state before another evaluation can run.

For applications that need native probability distributions over a defined set of answers, Vercel notes that Jev with AI SDK provides an alternative evaluation path. Jev's native probabilities and the OpenAI adapter's generated Boolean estimates come from different mechanisms, so any acceptance threshold should be tested against the specific provider and task.