Get it for macOS

Compare what every AI model costs per million tokens, across Anthropic, OpenAI, Fireworks AI, Ollama Cloud and OpenRouter, in one table. A Relative column ranks each model against the cheapest one on screen, weighted by an input/output blend you choose. Each row also gives the context window, the longest reply, and whether the model can reason, call tools or return structured output. Sort by any column, search by name, hide a provider with its legend dot, and open any model's own page from its row. Prices come from the public models.dev catalog and OpenRouter's model list and are kept for twelve hours; no account or key is needed, but the machine running the bb server must reach both sites.


Overview

View source

Choosing which model to point an agent at usually means opening five vendor pricing pages and doing arithmetic. The Model pricing page puts every model from the providers you follow in one table, with the numbers that decide the choice side by side.

What you get

  • A Model pricing page in the sidebar listing each model's price per million tokens in and out, its context window, its longest reply, and whether it can reason, call tools, return structured output and take a temperature.
  • A Relative column that ranks every model against the cheapest one on screen, so the cheapest reads 1.0× and the rest say how many times more they cost. How much of that blend is input is your choice: 80/20 suits an agent reading a repository, 25/75 a chat drafting prose.
  • Sorting by any column, search by name, and coloured provider dots that hide a provider's models for a moment without fetching anything.
  • Every row opens that model's own page on the provider's site.
  • A card per model in a narrow window, and palette commands for the page and its settings.

Where the prices come from

Anthropic, OpenAI and Fireworks AI publish no pricing API, so those three and Ollama Cloud are read from the community catalog at https://models.dev, and OpenRouter from its own public list. The bb server fetches them, keeps them for twelve hours, and refreshes on request; an unchanged catalog costs one round trip. If a source cannot be reached, its last prices stay on screen with a note saying how old they are.

Settings

  • One switch per provider: a provider switched off is neither fetched nor shown.
  • The input/output blend used by the Relative column.
  • Whether to list only models that can call tools (on by default). Models whose provider does not say are always shown.

What it needs

  • No account, key or login: only public price lists are read.
  • Network access from the machine running the bb server to models.dev and openrouter.ai.

Limits

  • Prices are as current as their source; the vendor's own pricing page remains the authority.
  • Only per-token prices are shown: no cached-input, batch, image or audio pricing.
  • With OpenRouter on, a model sold both directly and through the gateway appears twice, once per price.

More from Gustavo Ambrozio

FirstMate CrewHand one agent your to-do list and it runs a crew of worker threads, each in its own git worktree, until the work lands. The first mate is a pinned thread in bb's own chat that writes each worker's instructions, supervises it, and brings back branches, pull requests, findings and only the decisions that are yours. A FirstMate tab shows the crew as a board of Queued, Working, Blocked, Parked, Done, Failed and Idle, with Steer, Interrupt, Relaunch and End on every card. A built-in watch tells the first mate within about five minutes when a pull request is merged, reviewed or changes check status. Requires git, and the logged-in gh CLI for the pull request watch; a capable model such as Claude Sonnet is recommended for the first mate.NewGitHub BoardSee your GitHub issues, pull requests and discussions on one board, and hand any of them to an agent thread. Four columns cover what you wrote, what is open on repositories you own and what is assigned to you, and an issue that a pull request closes folds into that pull request's card. Pull request cards show CI checks and branch status, with an Update branch button wherever you can push. A detail panel shows the body, comments and images, and labels are added or removed from a right-click menu. Send to chat opens bb's new-thread composer on a card, with a prompt you set per column and per project, in the bb project that has the repository checked out. Requires the GitHub CLI (gh) installed and signed in on the machine running the bb server.1HeraldHear one spoken sentence when an agent needs you, and see every waiting thread in one list. When a thread asks a question, waits for a plan or a permission to be approved, finishes its turn or fails, the bb app says which thread it is and what it needs, on whatever page you are on. A Herald page in the sidebar lists every thread waiting on you with its reason and sentence, and the same sentence sits above that thread's composer with a Read again button until you answer. Settings choose which kinds of event are announced, the voice and speed, and where to speak: the desktop app, a browser tab, or the mobile app while it is open. By default the sentence is built from the event itself, with no model or external service; optionally, Claude Code, OpenAI Codex, Gemini CLI or a command of your own, installed and logged in on the Mac running bb, writes it instead, for one short model turn per announcement. The default voice needs bb running on a Mac, and elsewhere each device's own browser voice speaks.NewScheduled JobsRun shell commands on your Mac on a cron schedule or a fixed interval, as launchd jobs that keep running while bb is closed. Each job is a LaunchAgent plist that launchd owns, so a missed fire runs once when the Mac wakes and a disabled job stays disabled across reboots. A sidebar page shows each job's launchd status, its last twenty runs with duration and exit code, and the tail of its log, which you can follow live. Run now, Enable, Disable, Edit and Delete do the launchctl work for you, and commands run through your login shell with your terminal's PATH. The sidebar row counts failing jobs, and a picker switches between the Macs connected to bb. macOS only.New

More in Token Usage & Limits 20

Understand or control token, context-window, and provider-quota use.

View all