## How Pricing Works

ModelStack uses a **three-tier pricing model** — Economy, Standard, and Premium — with a wallet-based rate applied pay-as-you-go. There's no separate "within cap" price; the same rate table applies to a model every time you use it, whether or not you have a subscription. The twist: the rate isn't just a function of tier anymore — it depends on **which wallet pays**, since bonus and prepaid use different rate tables.

### Tier-Based Pricing

| Tier         | Bonus Wallet Rate                    | Prepaid Wallet Rate                          | Wallet Eligibility              |
| ------------ | -------------------------------------- | ----------------------------------------------- | ---------------------------------- |
| **Economy**  | 1.0x (no discount)                     | 0.2x by default (80% off provider cost)         | Bonus wallet or prepaid wallet     |
| **Standard** | 1.5x (50% markup)                      | 0.5x by default (50% off provider cost)         | Bonus wallet or prepaid wallet     |
| **Premium**  | Not eligible — bonus can't pay Premium | 1.0x (full provider cost, never discounted)     | Prepaid wallet only               |

<Info>
  Standard's prepaid rate can be lower for some model families: Anthropic
  defaults to 0.35x, OpenAI to 0.2x. Two specific models — `gpt-5.4` and
  `gpt-5.6-terra` — are pinned to 0.2x regardless of family. Economy's
  prepaid rate has no overrides yet — every model uses the 0.2x default.
  These overrides only ever touch the *prepaid* rate; bonus rates and
  Premium are never overridden — Premium is always exactly 1.0x.
</Info>

### Why Two Rates Per Tier?

Bonus balance and prepaid balance are priced differently on purpose. Bonus (referral/promo credit) is meant to give you free usage, not a discount, so on Economy it's worth exactly face value (1.0x) and on Standard it's actually worth *less* than face value (1.5x — a markup). Prepaid balance is your real money, so it gets the real discount: 0.2x on Economy, 0.5x on Standard by default. Premium has no bonus rate at all — it's prepaid-only, always 1.0x.

Because the two rates differ, a single Economy or Standard request can be billed from **both** wallets at once. Bonus covers as much of the request's raw provider cost as your available bonus balance affords at bonus's rate; prepaid covers whatever raw cost is left, at prepaid's rate. See [Example 4](#example-4-a-request-split-across-both-wallets) below for the full math.

### Which Tier Can a Model Use?

Models priced at $3/M input tokens or less are eligible for **all three tiers** — Economy and Standard are simply discounted ways to run them. Models priced above $3/M input are **Premium-only**. (This is the platform's default fallback rule; a model's actual eligible tiers can also be configured directly and may differ from this price cutoff.)

## Provider Pricing (Base Rates)

These are the upstream provider costs **before** any tier rate is applied. Your actual cost is lower on Economy or Standard, especially when paid from prepaid.

### Anthropic

| Model               | Available Tiers               | Input / 1M | Output / 1M |
| ------------------- | ------------------------------ | ---------- | ------------ |
| `claude-opus-4-6`   | Premium only                    | $5.00      | $25.00       |
| `claude-sonnet-4-6` | Economy + Standard + Premium    | $3.00      | $15.00       |
| `claude-sonnet-5`   | Economy + Standard + Premium    | $3.00      | $15.00       |
| `claude-haiku-4-5`  | Economy + Standard + Premium    | $1.00      | $5.00        |

### OpenAI

| Model          | Available Tiers               | Input / 1M | Output / 1M |
| -------------- | ------------------------------ | ---------- | ------------ |
| `gpt-5.5`      | Premium only                    | $5.00      | $30.00       |
| `gpt-5.4`      | Economy + Standard + Premium    | $2.50      | $15.00       |
| `gpt-4.1-mini` | Economy + Standard + Premium    | $0.40      | $1.60        |
| `gpt-4o-mini`  | Economy + Standard + Premium    | $0.15      | $0.60        |

### Google

| Model              | Available Tiers               | Input / 1M | Output / 1M |
| ------------------ | ------------------------------ | ---------- | ------------ |
| `gemini-3.1-pro`   | Economy + Standard + Premium    | $2.00      | $12.00       |
| `gemini-2.5-pro`   | Economy + Standard + Premium    | $1.25      | $10.00       |
| `gemini-2.5-flash` | Economy + Standard + Premium    | $0.30      | $2.50        |
| `gemma-4-31b-it`   | Economy + Standard + Premium    | $0.12      | $0.37        |

### DeepSeek

| Model               | Available Tiers               | Input / 1M | Output / 1M |
| ------------------- | ------------------------------ | ---------- | ------------ |
| `deepseek-v4-pro`   | Economy + Standard + Premium    | $0.44      | $0.87        |
| `deepseek-v3.2`     | Economy + Standard + Premium    | $0.25      | $0.38        |
| `deepseek-v4-flash` | Economy + Standard + Premium    | $0.11      | $0.22        |

### MiniMax

| Model          | Available Tiers               | Input / 1M | Output / 1M |
| -------------- | ------------------------------ | ---------- | ------------ |
| `minimax-m2.7` | Economy + Standard + Premium    | $0.28      | $1.20        |
| `minimax-m2.5` | Economy + Standard + Premium    | $0.15      | $1.15        |
| `minimax-m2.1` | Economy + Standard + Premium    | $0.29      | $0.95        |

### ZAI (GLM)

| Model     | Available Tiers               | Input / 1M | Output / 1M |
| --------- | ------------------------------ | ---------- | ------------ |
| `glm-5`   | Economy + Standard + Premium    | $0.60      | $1.92        |
| `glm-4.7` | Economy + Standard + Premium    | $0.40      | $1.75        |

### Qwen

| Model               | Available Tiers               | Input / 1M | Output / 1M |
| ------------------- | ------------------------------ | ---------- | ------------ |
| `qwen3-coder-next`  | Economy + Standard + Premium    | $0.11      | $0.80        |
| `qwen3-coder-flash` | Economy + Standard + Premium    | $0.20      | $0.98        |

## Cost Calculation Examples

### Example 1: Standard Model, Prepaid, Default Rate

**Request:** `gemini-2.5-flash` with 1,000 input tokens and 500 output tokens, Standard tier, paid entirely from prepaid

```
Provider cost:
├─ Input: 1,000 × $0.30 / 1,000,000 = $0.00030
├─ Output: 500 × $2.50 / 1,000,000 = $0.00125
└─ Total: $0.00155

Standard tier, prepaid rate (Google has no override, so the 0.5x default applies):
├─ Final cost: $0.00155 × 0.5 = $0.000775
└─ Paid from: prepaid wallet
```

### Example 2: Standard Model, Prepaid, Family Override

**Request:** `gpt-4o-mini` with 1,000 input tokens and 500 output tokens, Standard tier, paid entirely from prepaid

```
Provider cost:
├─ Input: 1,000 × $0.15 / 1,000,000 = $0.00015
├─ Output: 500 × $0.60 / 1,000,000 = $0.00030
└─ Total: $0.00045

Standard tier, prepaid rate (OpenAI family override: 0.2x instead of the 0.5x default):
├─ Final cost: $0.00045 × 0.2 = $0.00009
└─ Paid from: prepaid wallet
```

### Example 3: Premium Model

**Request:** `claude-opus-4-6` with 1,000 input tokens and 500 output tokens, Premium tier

```
Provider cost:
├─ Input: 1,000 × $5.00 / 1,000,000 = $0.005
├─ Output: 500 × $25.00 / 1,000,000 = $0.0125
└─ Total: $0.0175

Premium tier (always 1.0x, no exceptions):
├─ Final cost: $0.0175 × 1.0 = $0.0175
└─ Paid from: prepaid wallet only
```

### Example 4: A Request Split Across Both Wallets

**Request:** any Standard-tier model, $100 raw provider cost (scaled up for readable numbers), $30 of available bonus balance, prepaid wallet funded

```
Bonus wallet (rate 1.5x):
├─ Raw cost bonus can afford: $30 ÷ 1.5 = $20.00
└─ Charged to bonus: $30.00

Prepaid wallet (rate 0.5x, default):
├─ Remaining raw cost: $100 − $20 = $80.00
└─ Charged to prepaid: $80.00 × 0.5 = $40.00

Total charged: $30.00 + $40.00 = $70.00
Effective blended rate: $70.00 ÷ $100.00 = 0.7x
```

If this user had **no bonus balance at all**, the full $100 of raw cost would go to prepaid at 0.5x, costing **$50** — the plain default rate, no blending. The blended rate only shows up when a request draws from both wallets, and it depends on exactly how much bonus balance happened to be available at the time.

<Info>
  If prepaid has never been funded (balance at or below $0), the *entire*
  request — including what would normally be prepaid's portion — is billed
  from bonus at bonus's rate. An unfunded account is never charged real
  money.
</Info>

### Example 5: Same Model, All Three Tiers

**Request:** `claude-sonnet-4-6`, 1,000 input tokens and 500 output tokens — eligible for all three tiers since it's priced at $3/M input or less. Assume paid entirely from prepaid for a clean comparison.

```
Provider cost: $0.003 + $0.0075 = $0.0105

Economy (default prepaid rate, 0.2x):            $0.0105 × 0.2  = $0.00210
Standard (Anthropic family override, 0.35x):      $0.0105 × 0.35 = $0.003675
Premium (always 1.0x):                            $0.0105 × 1.0  = $0.0105
```

**Result:** Running the same model on Economy instead of Premium costs about a fifth as much here (when paid from prepaid); Standard lands in between.

## Cost Estimation Formula

```
provider_cost = (input_tokens × input_rate / 1,000,000) +
                (output_tokens × output_rate / 1,000,000)

# rate resolution depends on tier AND wallet:
# - Economy/Standard, prepaid: per-model override > per-family override > default (0.2 / 0.5)
# - Economy/Standard, bonus: fixed at 1.0 / 1.5, never overridden
# - Premium: always 1.0, prepaid only, never overridden
rate = tier_wallet_rate(model, tier, wallet)

final_cost = provider_cost × rate   # per wallet, then summed if split across wallets
```

## Pricing Best Practices

### 1. Prefer Economy or Standard When Possible

Both get a real discount when paid from prepaid — Economy at 0.2x by default, Standard at 0.5x by default (as low as 0.2x with a per-model or per-family override). Premium is always full price (1.0x), with no exceptions. Bonus balance doesn't get this discount: it's 1.0x on Economy and a 1.5x markup on Standard.

### 2. Check for Family and Model Overrides

Anthropic and OpenAI models default to a deeper Standard prepaid discount (0.35x and 0.2x respectively) than the general 0.5x default. `gpt-5.4` and `gpt-5.6-terra` are pinned to 0.2x specifically. Economy currently has no overrides — every model uses the 0.2x default. Other providers' Standard rate uses the 0.5x default.

### 3. Spend Bonus on Economy, Not Standard, When You Can

Since bonus is worth full face value on Economy (1.0x) but a 50% markup on Standard (1.5x), the same bonus balance stretches further on Economy-tier requests.

### 4. Reserve Premium for What Needs It

Premium is prepaid-wallet-only and never discounted — use it for models that are Premium-only (priced above $3/M input), or when you specifically need Premium routing for a model that's also available on Economy or Standard.

## Checking Costs in Dashboard

Your [Dashboard](https://modelstack.cc/dashboard/billing) shows:

- **Recent requests** with per-request costs
- **Cost breakdown** by model and tier
- **Wallet balance** and transaction history

<Info>
  All costs shown in the dashboard are **after** the tier and wallet rate are
  applied — the actual amount charged to your wallet(s). For a request split
  across bonus and prepaid, the shown rate is the blended effective rate
  described in [Example 4](#example-4-a-request-split-across-both-wallets)
  above, not a single table value.
</Info>
