OpenAI-compatible API
Milly Lab API
The Milly Lab API is an OpenAI-compatible REST API at https://api.m-lab.app/v1. One key reaches 13 chat models from OpenAI, Anthropic, Google, DeepSeek and xAI, 4 image models and an embeddings model. It comes with the Max plan ($42 a month) and higher. A call costs the provider's rate plus $1 per million input or output tokens.
Last updated:
Milly Lab API at a glance
| Item | Value |
|---|---|
| Base URL | https://api.m-lab.app/v1 |
| Protocol | OpenAI wire format for chat completions, embeddings and images; streaming over server-sent events |
| Authentication | Bearer key that starts with mk_, shown once when you create it |
| Plans with API access | Max, Advanced and Enterprise |
| Price | Provider rate + $1 per 1M input tokens and $1 per 1M output tokens; images + $0.01 each |
| Beyond the plan allowance | Paid from the top-up balance at 2× the listed price |
| Rate limit | Up to 120 requests per minute per key |
| Keys per account | Up to 5 active keys, each with its own spend cap and model list |
| Live price list | GET /v1/models, no key required |
How do I make my first request?
Create a key under API keys in your Milly Lab account, then point any OpenAI SDK at https://api.m-lab.app/v1. Each example below sends the same chat request.
- Create a key. It starts with mk_ and is shown only once, so store it in an environment variable such as MILLY_API_KEY.
- Set the base URL to https://api.m-lab.app/v1. Nothing else changes in OpenAI-compatible code.
- Pick a model id from GET /v1/models, for example claude-sonnet-5.
curl
curl https://api.m-lab.app/v1/chat/completions \
-H "Authorization: Bearer $MILLY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "claude-sonnet-5", "messages": [{"role": "user", "content": "Hello"}]}'Python
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.m-lab.app/v1", api_key=os.environ["MILLY_API_KEY"])
r = client.chat.completions.create(
model="claude-sonnet-5",
messages=[{"role": "user", "content": "Hello"}],
)
print(r.choices[0].message.content)Node.js
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.m-lab.app/v1", apiKey: process.env.MILLY_API_KEY });
const r = await client.chat.completions.create({
model: "claude-sonnet-5",
messages: [{ role: "user", content: "Hello" }],
});
console.log(r.choices[0].message.content);The response is a standard chat.completion object, and its usage field shows the tokens you were billed for. Send stream: true to receive the answer as server-sent events instead.
Which endpoints does the API have?
All 5 API endpoints live under https://api.m-lab.app/v1.
| Endpoint | What it does |
|---|---|
GET /v1/models | Lists every API model with live prices, context window and output limit. Works without a key. |
POST /v1/chat/completions | Chat completion, streamed over server-sent events or returned in one response. |
POST /v1/embeddings | Embeddings with text-embedding-3-small (1,536 dimensions). |
POST /v1/images/generations | Generates 1–4 images per call with 4 image models. Synchronous, up to 180 seconds. |
GET /v1/usage | Requests, tokens and billed dollars per key and per model, for up to 92 days per query. |
Which plans include API access?
API access is included in the Max, Advanced and Enterprise plans. On Free, Starter and Pro you cannot create API keys. If your plan drops below Max, existing keys stop working until you upgrade, but you can still see and revoke them.
| Plan | Price per month | API access |
|---|---|---|
| Free | $0 | No |
| Starter | $12 | No |
| Pro | $25 | No |
| Max | $42 | Yes |
| Advanced | $72 | Yes |
| Enterprise | Custom | Yes |
The Free plan gives 10 chat messages per model every 4 hours in the web app, without API access.
How much does an API call cost?
An API call costs the provider's rate for the model plus a Milly Lab platform fee of $1 per million input tokens and $1 per million output tokens; image generation adds $0.01 per image. Your plan's monthly allowance for the model pays first, and anything beyond it comes from your top-up balance at 2× the price.
- Plan allowance first: a model your plan allocates is paid from that model's monthly allowance at the listed price.
- Then the top-up balance at 2×: other models, and usage after the allowance runs out, cost 2× the listed price.
- Checked before the call: if allowance and balance cannot cover a request, the API answers 402 before any provider is called, so nothing is billed.
- Partial output is billed: if your client disconnects mid-stream, generation finishes and the tokens produced are charged.
Worked example
A Claude Sonnet 5 request with 10,000 input tokens and 1,000 output tokens costs $0.041 from the plan allowance (10,000 × $3.00 + 1,000 × $11.00 per million tokens) or $0.082 from the top-up balance.
Prices are in US dollars. Payments are processed by SmartPay, a payment system from Tajikistan, and can be made in any currency.
Which models can I call, and at what price?
These are the models GET /v1/models listed on 2026-10-04: chat models, image models and one embeddings model. Prices include the $1 platform fee; the endpoint itself is always the current source.
Chat models
| Model | API id | Input, $ per 1M tokens | Output, $ per 1M tokens | Context window, tokens | Plan allowance |
|---|---|---|---|---|---|
| GPT-5.6 Luna | gpt-5.4-nano | $1.20 | $2.20 | 400,000 | Max, Advanced |
| GPT-5.6 Terra | gpt-5.4 | $3.00 | $13.00 | 400,000 | Max, Advanced |
| GPT-5.6 Sol | gpt-5.5 | $5.00 | $21.00 | 400,000 | Max, Advanced |
| ChatGPT Codex | gpt-5.3-codex | $2.75 | $15.00 | 400,000 | Advanced |
| Claude Haiku 4.5 | claude-haiku-4-5 | $2.00 | $6.00 | 200,000 | Max, Advanced |
| Claude Sonnet 5 | claude-sonnet-5 | $3.00 | $11.00 | 1,000,000 | Max, Advanced |
| Claude Opus 5 | claude-opus-4-8 | $6.00 | $26.00 | 200,000 | Advanced |
| DeepSeek V4 Flash | deepseek-v4-flash | $1.22 | $1.66 | 128,000 | Balance only |
| DeepSeek V4 | deepseek-v4 | $1.22 | $1.66 | 128,000 | Balance only |
| DeepSeek V4 Pro | deepseek-v4-pro | $1.66 | $2.98 | 256,000 | Max, Advanced |
| Gemini 3.7 Flash 200K | gemini-3.6-flash-200k | $1.75 | $4.75 | 200,000 | Balance only |
| Gemini 3.7 Flash | gemini-3.6-flash | $1.75 | $4.75 | 1,000,000 | Max, Advanced |
| Grok 4.6 | grok-4.3 | $3.00 | $7.00 | 256,000 | Balance only |
Image models
| Model | API id | Price per image, standard quality | Plan allowance |
|---|---|---|---|
| Gemini Nano Banana 2 | gemini-nano-banana | $0.077 | Max, Advanced |
| ChatGPT Image 2.0 | gpt-image-2 | $0.0627 | Max, Advanced |
| Flux 2 Pro | fal-ai/flux-pro | $0.04 | Max |
| Stable Diffusion 3.5 Large | fal-ai/stable-diffusion-xl | $0.075 | Advanced |
Embeddings
text-embedding-3-small returns 1,536-dimension vectors and costs $1.02 per million input tokens including the fee. No plan includes embeddings, so they are always paid from the top-up balance at 2×.
The Plan allowance column names the plans that include a monthly allowance for the model. Models marked “Balance only”, such as Grok 4.6, have no plan allowance and are paid from the top-up balance at 2×. Enterprise allowances are agreed per contract. Other image quality levels change the per-image price.
What are the rate limits?
Each API key can make up to 120 requests per minute, and an account can hold up to 5 active keys. A request over the limit gets HTTP 429 with a Retry-After header that says how many seconds to wait.
| Limit | Value |
|---|---|
| Requests per key | Up to 120 per minute |
| Active keys per account | 5 |
| Messages per chat request | 200 |
| Characters per chat request | 400,000 |
| Embedding inputs per call | 256, up to 12,000 characters each |
| Images per call | 1–4, prompt up to 4,000 characters |
| Time limit per image request | 180 seconds |
| Usage report range | Up to 92 days |
Every /v1 response carries an x-request-id header; quote it when you contact support. Responses to a request with a valid key also carry x-ratelimit-limit, x-ratelimit-remaining and x-ratelimit-window.
How do spend caps and allowed models work?
Each key can have a monthly spend cap in US dollars and a list of allowed models. A call after the cap is reached returns 402 spend_cap_reached; a call to a model outside the list returns 403 model_not_allowed. Both checks run before anything is charged.
| Control | How it works |
|---|---|
| Monthly spend cap | Whole cents, up to $100,000. It counts the billed spend on that key since the 1st of the month (UTC) and resets on the 1st. The request that crosses the cap still completes, so the overshoot is at most one request. |
| Allowed models | Up to 64 model ids; an empty list means every API model. Older aliases are resolved to the current id. |
| Revoking | A revoked key fails from its next request. You can list and revoke keys on any plan. |
CORS is open on /v1, so browser code can call the API directly, but a key in a web page is visible to every visitor. Give such a key a tight spend cap and an allowed-model list, or call the API from your own server.
What error codes does the API return?
Errors use the OpenAI error format shown below, so SDK error handling works unchanged. Branch on the code field, not on the message text.
{"error": {"message": "…", "type": "…", "code": "…"}}| HTTP | code | Meaning |
|---|---|---|
| 400 | invalid_body · invalid_messages · invalid_input · input_too_long · unsupported_encoding · invalid_range | Malformed body, no user text, invalid embedding input or an invalid usage date range |
| 401 | missing_api_key · invalid_api_key | No bearer key, or an unknown or revoked key |
| 402 | insufficient_for_request · insufficient_balance | Plan allowance plus balance cannot cover the request |
| 402 | spend_cap_reached | The key reached its monthly spend cap |
| 403 | api_access_required | The account's plan has no API access |
| 403 | model_not_allowed | The model is not on the key's allowed list |
| 403 | account_disabled | The account is disabled |
| 404 | model_not_found · key_not_found | Unknown or retired model, or a key id that is not on this account |
| 429 | rate_limit_exceeded | More than 120 requests per minute on this key; wait for the Retry-After time |
| 502 | upstream_error | The provider failed before producing output; nothing is billed |
| 503 | provider_unavailable | The model's provider is currently unavailable; nothing is billed |
Questions developers ask
Does the Milly Lab API work with the OpenAI SDKs?
Yes. Chat completions, embeddings and image generation use the OpenAI wire format, so the official openai packages for Python and Node.js, and other OpenAI-compatible tools, work once you set the base URL to https://api.m-lab.app/v1 and use a Milly Lab key.
Does the API support tool calling and image input?
Not yet. The tools and tool_choice fields are accepted so SDK defaults do not fail, but they are not executed. Chat messages take text only; image parts are ignored.
Can I generate video or 3D through the API?
No. The API covers chat completions, embeddings and image generation. Video and 3D generation are not available through /v1.
Can I call a model my plan does not include?
Yes. On the Max plan and higher you can call any model that GET /v1/models lists. A model without a monthly allowance on your plan, such as Grok 4.6, is paid from your top-up balance at 2× the listed price.
Is there a free API tier?
No. Creating API keys requires the Max plan or higher. The Free plan gives 10 chat messages per model every 4 hours in the web app, without API access.
How do I pay for API usage?
API usage is paid from your plan's monthly allowance first and then from your top-up balance; prices are in US dollars. Payments are processed by SmartPay, a payment system from Tajikistan, and can be made in any currency.
Where is the full API reference?
The API guide covers streaming, embeddings, images, the usage endpoint and best practices in detail.