Intents
macOS · Developer toolOpen source on GitHub

Intents

A closer look at what your model actually does.

By Cory Parry · Updated

View on GitHub
Intents app icon on its turquoise and coral banner
Intents: a native macOS workbench for Apple Foundation Models.

About the project

Intents is a native macOS evaluation workbench for Apple’s Foundation Models. It gives prompts, test cases, scoring, and run history a shared place, so you can investigate model behaviour and compare changes with a saved baseline.

Build repeatable evaluations

Organise cases, instructions, and scoring into suites you can run again as you change your prompts or model settings.

Look beyond the answer

Inspect responses, scores, timing, and execution traces to understand how a run reached its result.

Compare with a baseline

Keep run history and compare saved evaluations to see where behaviour improves or regresses.

Who is Intents for?

Intents is for developers testing prompts, instructions, and tools built around Apple’s Foundation Models. Saved suites and baseline comparisons make it useful when you need to understand whether a change improves the responses your app depends on.

Install Intents

  1. Download the latest macOS DMG from the GitHub releases page.
  2. Open the disk image, drag Intents into Applications, then eject the disk image.
  3. Launch the app and check that the on-device model is ready. Xcode is needed to build from source or use Intent Lab; normal on-device evaluations do not require it.
Download the latest release ↗

Run your first evaluation

  1. Create a suite in Overview, then open Cases to add prompts and expected responses or reference answers.
  2. Use Setup → Instructions for shared instructions and reference files, then Setup → Scoring to choose a scoring method and repetitions.
  3. Run the suite and open its Workflow trace or Report to inspect responses, scores, explanations and tool activity.
  4. Open saved runs from Results or the sidebar, then use Compare to compare earlier runs or export a JSON report.

Requirements and limitations

  • macOS 27 or later and a Mac that supports Apple Intelligence.
  • For the default on-device provider, enable Apple Intelligence and download its model.
  • Xcode 27 is required to build from source or use Intent Lab; it is not required for normal on-device evaluations.

An evaluation measures the cases you provide. Review AI-rubric explanations and use verified reference answers for factual tasks; a few repetitions do not establish statistical significance. On-device inference runs locally, while configured HTTP providers and tools receive the content needed for their calls. Exported traces can contain prompts, responses, and tool data.

Read the setup, provider, and data-handling guide ↗

Evaluating Apple AI models: common questions

How can I evaluate Apple’s Foundation Models on my Mac?

Use Intents to create a suite of representative prompts, shared instructions and reference answers. Choose a scoring mode, run repeated trials, and inspect the responses and execution traces. Save the run as a baseline before changing your prompt or model settings. The default provider uses Apple’s on-device Foundation Model and requires a supported Mac with Apple Intelligence enabled and its model downloaded.

Which scoring method should I use for an Apple AI evaluation?

Use exact-text scoring when the entire expected response matters, contains-text scoring for required phrases, and an AI rubric for explicit quality criteria. Collect-only mode records responses and traces without assigning a score. Intents’ AI rubric uses a 1–4 scale, with 3 or 4 passing. Read the explanations and check reference answers: a rubric score is evidence for those test cases, not a general measure of model intelligence.

How do I compare model results and spot regressions?

Run the same suite before and after a change, save both runs, and choose the earlier run as the baseline in Results and Compare. Inspect case-level responses, scores, timing and tool activity rather than relying only on an overall average. Keep the cases, scoring criteria and settings comparable, repeat variable tasks, and record your Mac and macOS version when sharing results.

Can Codex or another MCP-compatible agent run evaluations?

Yes. Intents includes an MCP server for managing suites and references, running evaluations, inspecting traces and comparing saved results. For Codex, open Settings, choose Connect to Codex, restart Codex and keep Intents open. Other compatible clients can use the documented local HTTP endpoint at http://127.0.0.1:17873/mcp with the managed bearer credential. Intents stores this credential in the login Keychain and configures Codex to send it. The connector listens on the same Mac; authenticated local clients can read and change evaluation data.

Where can I view or share Intents results?

Review individual responses and traces in the app’s Results, compare saved runs, or export a JSON report. This page describes the evaluation tool; it does not publish a model leaderboard or claim benchmark scores. Before sharing an export, review it for prompts, responses, attachments and tool data. A useful report includes the suite, scoring criteria, model/provider, settings, repetitions and test environment.

Do Apple model evaluations run entirely on-device?

The default on-device provider runs inference locally. Configured HTTP providers and tools receive the data needed for their calls, and optional Private Cloud Compute uses Apple’s network service when available. Saved runs and exported traces can contain sensitive content. Review the project’s data-handling guide before evaluating private documents or sharing reports.

Sources and further reading

Plain-text project guides

Written by Cory Parry, the developer of Intents. Release details and requirements can change; use the linked project documentation for the latest information.