Newt Brings On-Device Multiple-Choice Questioning to Swift Developers

A new Swift package called Newt is making it possible to extract structured, multiple-choice answers from text using a language model that runs entirely on a user's machine. Developed by willswire, the package targets a specific but growing need: applications that need to classify or route text into predefined categories without sending data to a remote server.

Newt takes a piece of text and answers questions about it, supporting three formats: multiple-choice with two to 255 options, graded scoring across two to ten levels, and yes/no queries the package calls "Noul." For each option, it returns a probability, a confidence score, and a coverage metric that indicates whether the model actually answered within the expected format.

How It Works Under the Hood

Newt does not ship or redistribute model weights. Instead, users export an open-source model through Apple's coreai-models toolchain, which compiles the model for execution on Apple Silicon via Core AI. The package currently supports Qwen3-4B, which requires about 2.1 GB of disk space. A smaller Qwen3-0.6B variant also runs but fails roughly half of the test suite's fixture cases, so the 4B model is the recommended starting point.

The scoring mechanism sums the log-probability of each option's tokens plus the end-of-turn token after a fixed prompt, then normalizes those values across all options. This mirrors the standard multiple-choice evaluation approach used in broader NLP research, but wrapped in a Swift-native API that mirrors the conventions of TypeSafe, a similar tool for evaluating language models.

Performance and Constraints

On an Apple M3 Pro with 18 GB of memory running macOS 27.0, the first model load takes between 7.6 and 11.8 seconds because Core AI specializes the model for the specific hardware and caches the result. Subsequent loads drop to approximately one second. A single question with three options takes about half a second on the first ask after loading, and roughly 435 milliseconds once warm. Two Noul questions together take around 600 milliseconds.

Debug builds are significantly slower: 4.5 times slower on the 4B model and 14 times slower on the 0.6B model. The package must be measured with release builds for any meaningful performance data.

There are limitations. The package has been tested only on Apple Silicon Macs with macOS 27 and Xcode 27; Core AI does not exist on Linux. iOS support is unproven on actual devices. Models that introduce a BOS token in their chat template can double the prompt length. And because Newt reuses the model's cache assuming a pure transformer architecture, it does not work with hybrid or state-space models.

Design Decisions and Trade-Offs

Newt deliberately mirrors TypeSafe's naming conventions and API shape. Terms like "state," "criteria," "Choice," "Score," and "Noul" carry over directly, meaning examples from TypeSafe's documentation work in Newt with minimal changes. The package is not affiliated with TypeSafe but is clearly built to fill a complementary role: where TypeSafe evaluates models, Newt applies them.

The probability distributions returned by Newt should be treated as rankings rather than calibrated confidence levels. The package's author is explicit about this: on ambiguous inputs, wording can dramatically flip results. In one test case asking whether a customer wanted a human agent, the probability shifted from 0.9996 to 0.0000004 depending on prompt phrasing, while TypeSafe returned 0.84.

The confidence metric, calculated as (max p − 1/n) / (1 − 1/n), reaches near 1.0 across every fixture test on the 4B model, making it less informative in practice than the formula suggests. Coverage, which measures the share of probability mass assigned to listed options, is more useful: low coverage signals that the model wanted to say something outside the provided choices.

Practical Implications

For developers building iOS or macOS applications that need to classify text without routing data through an API, Newt offers a working path. The full test suite passes with outbound network access blocked, confirming the on-device claim. The package is MIT-licensed, and its entire codebase was written with the assistance of Claude Code, Anthropic's coding agent.

It is a niche tool with clear boundaries: a single question at a time, a limited set of tested models, and a platform that will not expand beyond Apple Silicon anytime soon. But for the developers who fit that profile, it removes the privacy and latency trade-off that typically comes with running language model inference on user data.