Начало работы

Plans, Billing & Usage

This guide explains what your plan includes, how the monthly allowance is spent, what happens when it runs out, and how to read the usage view. It is for anyone on a paid or free Milly Lab plan.


#How Milly Lab charges for model use

Every model costs a different amount to run. Rather than selling you a message count, Milly Lab gives each plan a monthly allowance in US dollars, divided across a specific set of models. Each model has its own portion of that allowance.

Two consequences:

  1. You can exhaust one model while others still have budget. Running out of your Claude Sonnet allowance does not stop you using GPT.
  2. Plans are not cumulative. A plan includes exactly the models listed for it — not the models of the plans below it. Moving up a plan can mean losing access to a model, so check the model list, not just the price.

#What each plan includes

Free

No monetary allowance. A registered user without a subscription gets chat only, on two models:

Modelsgpt-5.4-nano, deepseek-v4-flash
Limit10 messages per model per 4-hour window
Everything elseLocked — image, video, 3D, Swarm Mind, file generation

The window is rolling. When you hit the limit, the error tells you when it resets.

Starter — $7.00 per month

ModelAllowanceApprox. messages
gpt-5.4-nano$1.251,445
claude-haiku-4-5$3.60973
deepseek-v4-flash$1.404,545
gemini-2.5-flash$0.75466

Chat only. No file generation, no Swarm Mind.

Pro — $15.00 per month

ModelAllowanceApprox. output
gpt-5.4$4.00281 messages
claude-sonnet-5$4.50300 messages
deepseek-v4$2.001,576 messages
gemini-3.6-flash-200k$2.25273 messages
gemini-nano-banana$2.2557 images

Adds file generation. No Swarm Mind.

Max — $27.00 per month

ModelAllowanceApprox. output
gpt-5.5$6.25156 messages
claude-sonnet-5$4.65310 messages
deepseek-v4-pro$2.50359 messages
gemini-3.6-flash$2.10180 messages
gemini-nano-banana$3.2583 images
gpt-image-2$2.0037 images
fal-ai/flux-pro$1.2021 images
veo-3.1$5.055 clips

Adds Swarm Mind and video generation.

Advanced — $52.00 per month

ModelAllowanceApprox. output
gpt-5.5$8.75219 messages
claude-sonnet-5$5.40360 messages
claude-opus-4-8$6.75193 messages
deepseek-v4-pro$3.75539 messages
gemini-3.6-flash$3.50300 messages
gpt-5.3-codex$7.50226 sessions
gemini-nano-banana$3.0578 images
gpt-image-2$1.6030 images
fal-ai/stable-diffusion-xl$0.60100 images
veo-3.1$5.005 clips
gen4_turbo$3.504 clips
kling-3.0$2.004 clips
meshy-api$0.603 models

Everything, including 3D generation.

Enterprise

Enterprise plans are arranged directly — get in touch through the Contacts page.

The message counts are estimates, calculated from a reference message of roughly 4,000 input tokens and 750 output tokens. A long conversation, a large attachment, or a run that reads files from a connected repository consumes more per message. Treat them as a guide, not a quota.


#What uses your allowance

Anything that calls a model:

  • Chat messages, including the conversation history sent as context each turn.
  • Image, video and 3D generation.
  • Every model call inside a Swarm Mind run — an orchestrator plus several workers means several charges from one brief.
  • File contents read from a connected GitHub repository or Google Drive, which are sent as input tokens.
  • Web search, when enabled for a message.

Swarm Mind is the expensive one. A Project Mode run may call models many times. A run over a connected repository adds the file contents on top.

Web search adds cost and latency. It is a toggle in the composer — leave it off unless the answer genuinely depends on current information.


#When your allowance runs out

Your prepaid balance takes over, at twice the underlying provider cost. The markup is why the balance is a fallback rather than the main way to pay — a plan is the cheaper route for regular use.

If both the model's allowance and your balance are insufficient, the request is refused before it runs. You are not charged for a partial answer, and nothing is billed silently.

Milly Lab also checks in advance that a request can pay for its full input and a usable answer. This is why you may be blocked from a request that looks affordable — a long conversation with a large attachment can cost more than the balance remaining, and Milly Lab refuses rather than cutting the answer off mid-sentence.


#Reading the errors

All of these arrive as HTTP 402 Payment Required with a machine-readable code.

CodeWhat it meansWhat to do
free_limitFree tier: 10 messages used for this model in the current 4-hour windowWait for the reset time in the message, use the other free model, or subscribe
free_model_lockedFree tier: this model needs a subscriptionUse gpt-5.4-nano or deepseek-v4-flash, or subscribe
plan_requiredThe feature needs any paid plan; the free tier cannot buy into it with balanceSubscribe
feature_requiredYour plan lacks this capability — file generation needs Pro or aboveMove to a plan that includes it
insufficient_balanceThe model's allowance is spent and your balance cannot cover the next callTop up, or use a model that still has allowance
insufficient_for_requestThis specific request is too large for what remainsShorten the conversation, remove attachments, start a new conversation, or top up

insufficient_for_request catches people out. It usually means the conversation has grown, not that you are out of budget generally. Starting a fresh conversation often resolves it, because the history is no longer being resent.


#Where to see your usage

Go to Settings → Billing (/app/settings/billing).

It shows your current plan, how much of each model's allowance is left, your balance, and your transaction history.

Every billable turn is recorded with the model, input tokens, output tokens, cost, and whether it came from your plan allowance or your balance. If a charge looks wrong, that record is what to check.


#Billing periods

Allowances are granted per subscription period and reset at the start of each one. Unused allowance does not carry over.

When a period ends without renewal, access falls back to the free tier. Your conversations, generated files and connected workspaces are not deleted — you keep everything, you just lose access to the paid models until you resubscribe.


#Paying

No payment gateway is connected yet. Milly Lab cannot currently take payment. A gateway with local banks is planned.

Until then the upgrade and top-up screens will not complete a purchase.

To arrange a plan in the meantime, get in touch through the Contacts page.


#Common questions

Can I move allowance between models? No. Each model's portion is fixed for the plan.

Why did a Swarm Mind run cost so much more than a chat message? It made many model calls — an orchestrator and several workers, each with its own input and output. Connected-workspace file contents add to that.

Why does the same prompt cost different amounts? The whole conversation is sent as context each turn, so cost grows as a conversation lengthens. Attachments and connected files add to it.

I upgraded and lost a model. Plans are not cumulative. Check the model list for the plan you moved to.

Does my team share my allowance? Every billable turn is recorded per user in the usage view. For how your plan meters team usage, ask through the Contacts page.