OpenAI-compatible API

Milly Lab API

The Milly Lab API is an OpenAI-compatible REST API at https://api.m-lab.app/v1. One key reaches 13 chat models from OpenAI, Anthropic, Google, DeepSeek and xAI, 4 image models and an embeddings model. It comes with the Max plan ($42 a month) and higher. A call costs the provider's rate plus $1 per million input or output tokens.

Last updated:

Milly Lab API at a glance

ItemValue
Base URLhttps://api.m-lab.app/v1
ProtocolOpenAI wire format for chat completions, embeddings and images; streaming over server-sent events
AuthenticationBearer key that starts with mk_, shown once when you create it
Plans with API accessMax, Advanced and Enterprise
PriceProvider rate + $1 per 1M input tokens and $1 per 1M output tokens; images + $0.01 each
Beyond the plan allowancePaid from the top-up balance at 2× the listed price
Rate limitUp to 120 requests per minute per key
Keys per accountUp to 5 active keys, each with its own spend cap and model list
Live price listGET /v1/models, no key required

How do I make my first request?

Create a key under API keys in your Milly Lab account, then point any OpenAI SDK at https://api.m-lab.app/v1. Each example below sends the same chat request.

  1. Create a key. It starts with mk_ and is shown only once, so store it in an environment variable such as MILLY_API_KEY.
  2. Set the base URL to https://api.m-lab.app/v1. Nothing else changes in OpenAI-compatible code.
  3. Pick a model id from GET /v1/models, for example claude-sonnet-5.

curl

curl https://api.m-lab.app/v1/chat/completions \
  -H "Authorization: Bearer $MILLY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "claude-sonnet-5", "messages": [{"role": "user", "content": "Hello"}]}'

Python

import os
from openai import OpenAI

client = OpenAI(base_url="https://api.m-lab.app/v1", api_key=os.environ["MILLY_API_KEY"])
r = client.chat.completions.create(
    model="claude-sonnet-5",
    messages=[{"role": "user", "content": "Hello"}],
)
print(r.choices[0].message.content)

Node.js

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.m-lab.app/v1", apiKey: process.env.MILLY_API_KEY });
const r = await client.chat.completions.create({
  model: "claude-sonnet-5",
  messages: [{ role: "user", content: "Hello" }],
});
console.log(r.choices[0].message.content);

The response is a standard chat.completion object, and its usage field shows the tokens you were billed for. Send stream: true to receive the answer as server-sent events instead.

Which endpoints does the API have?

All 5 API endpoints live under https://api.m-lab.app/v1.

EndpointWhat it does
GET /v1/modelsLists every API model with live prices, context window and output limit. Works without a key.
POST /v1/chat/completionsChat completion, streamed over server-sent events or returned in one response.
POST /v1/embeddingsEmbeddings with text-embedding-3-small (1,536 dimensions).
POST /v1/images/generationsGenerates 1–4 images per call with 4 image models. Synchronous, up to 180 seconds.
GET /v1/usageRequests, tokens and billed dollars per key and per model, for up to 92 days per query.

Which plans include API access?

API access is included in the Max, Advanced and Enterprise plans. On Free, Starter and Pro you cannot create API keys. If your plan drops below Max, existing keys stop working until you upgrade, but you can still see and revoke them.

PlanPrice per monthAPI access
Free$0No
Starter$12No
Pro$25No
Max$42Yes
Advanced$72Yes
EnterpriseCustomYes

The Free plan gives 10 chat messages per model every 4 hours in the web app, without API access.

Compare all plans

How much does an API call cost?

An API call costs the provider's rate for the model plus a Milly Lab platform fee of $1 per million input tokens and $1 per million output tokens; image generation adds $0.01 per image. Your plan's monthly allowance for the model pays first, and anything beyond it comes from your top-up balance at 2× the price.

  • Plan allowance first: a model your plan allocates is paid from that model's monthly allowance at the listed price.
  • Then the top-up balance at 2×: other models, and usage after the allowance runs out, cost 2× the listed price.
  • Checked before the call: if allowance and balance cannot cover a request, the API answers 402 before any provider is called, so nothing is billed.
  • Partial output is billed: if your client disconnects mid-stream, generation finishes and the tokens produced are charged.

Worked example

A Claude Sonnet 5 request with 10,000 input tokens and 1,000 output tokens costs $0.041 from the plan allowance (10,000 × $3.00 + 1,000 × $11.00 per million tokens) or $0.082 from the top-up balance.

Prices are in US dollars. Payments are processed by SmartPay, a payment system from Tajikistan, and can be made in any currency.

Which models can I call, and at what price?

These are the models GET /v1/models listed on 2026-10-04: chat models, image models and one embeddings model. Prices include the $1 platform fee; the endpoint itself is always the current source.

Chat models

ModelAPI idInput, $ per 1M tokensOutput, $ per 1M tokensContext window, tokensPlan allowance
GPT-5.6 Lunagpt-5.4-nano$1.20$2.20400,000Max, Advanced
GPT-5.6 Terragpt-5.4$3.00$13.00400,000Max, Advanced
GPT-5.6 Solgpt-5.5$5.00$21.00400,000Max, Advanced
ChatGPT Codexgpt-5.3-codex$2.75$15.00400,000Advanced
Claude Haiku 4.5claude-haiku-4-5$2.00$6.00200,000Max, Advanced
Claude Sonnet 5claude-sonnet-5$3.00$11.001,000,000Max, Advanced
Claude Opus 5claude-opus-4-8$6.00$26.00200,000Advanced
DeepSeek V4 Flashdeepseek-v4-flash$1.22$1.66128,000Balance only
DeepSeek V4deepseek-v4$1.22$1.66128,000Balance only
DeepSeek V4 Prodeepseek-v4-pro$1.66$2.98256,000Max, Advanced
Gemini 3.7 Flash 200Kgemini-3.6-flash-200k$1.75$4.75200,000Balance only
Gemini 3.7 Flashgemini-3.6-flash$1.75$4.751,000,000Max, Advanced
Grok 4.6grok-4.3$3.00$7.00256,000Balance only

Image models

ModelAPI idPrice per image, standard qualityPlan allowance
Gemini Nano Banana 2gemini-nano-banana$0.077Max, Advanced
ChatGPT Image 2.0gpt-image-2$0.0627Max, Advanced
Flux 2 Profal-ai/flux-pro$0.04Max
Stable Diffusion 3.5 Largefal-ai/stable-diffusion-xl$0.075Advanced

Embeddings

text-embedding-3-small returns 1,536-dimension vectors and costs $1.02 per million input tokens including the fee. No plan includes embeddings, so they are always paid from the top-up balance at 2×.

The Plan allowance column names the plans that include a monthly allowance for the model. Models marked “Balance only”, such as Grok 4.6, have no plan allowance and are paid from the top-up balance at 2×. Enterprise allowances are agreed per contract. Other image quality levels change the per-image price.

What are the rate limits?

Each API key can make up to 120 requests per minute, and an account can hold up to 5 active keys. A request over the limit gets HTTP 429 with a Retry-After header that says how many seconds to wait.

LimitValue
Requests per keyUp to 120 per minute
Active keys per account5
Messages per chat request200
Characters per chat request400,000
Embedding inputs per call256, up to 12,000 characters each
Images per call1–4, prompt up to 4,000 characters
Time limit per image request180 seconds
Usage report rangeUp to 92 days

Every /v1 response carries an x-request-id header; quote it when you contact support. Responses to a request with a valid key also carry x-ratelimit-limit, x-ratelimit-remaining and x-ratelimit-window.

How do spend caps and allowed models work?

Each key can have a monthly spend cap in US dollars and a list of allowed models. A call after the cap is reached returns 402 spend_cap_reached; a call to a model outside the list returns 403 model_not_allowed. Both checks run before anything is charged.

ControlHow it works
Monthly spend capWhole cents, up to $100,000. It counts the billed spend on that key since the 1st of the month (UTC) and resets on the 1st. The request that crosses the cap still completes, so the overshoot is at most one request.
Allowed modelsUp to 64 model ids; an empty list means every API model. Older aliases are resolved to the current id.
RevokingA revoked key fails from its next request. You can list and revoke keys on any plan.

CORS is open on /v1, so browser code can call the API directly, but a key in a web page is visible to every visitor. Give such a key a tight spend cap and an allowed-model list, or call the API from your own server.

What error codes does the API return?

Errors use the OpenAI error format shown below, so SDK error handling works unchanged. Branch on the code field, not on the message text.

{"error": {"message": "…", "type": "…", "code": "…"}}
HTTPcodeMeaning
400invalid_body · invalid_messages · invalid_input · input_too_long · unsupported_encoding · invalid_rangeMalformed body, no user text, invalid embedding input or an invalid usage date range
401missing_api_key · invalid_api_keyNo bearer key, or an unknown or revoked key
402insufficient_for_request · insufficient_balancePlan allowance plus balance cannot cover the request
402spend_cap_reachedThe key reached its monthly spend cap
403api_access_requiredThe account's plan has no API access
403model_not_allowedThe model is not on the key's allowed list
403account_disabledThe account is disabled
404model_not_found · key_not_foundUnknown or retired model, or a key id that is not on this account
429rate_limit_exceededMore than 120 requests per minute on this key; wait for the Retry-After time
502upstream_errorThe provider failed before producing output; nothing is billed
503provider_unavailableThe model's provider is currently unavailable; nothing is billed

Questions developers ask

Does the Milly Lab API work with the OpenAI SDKs?

Yes. Chat completions, embeddings and image generation use the OpenAI wire format, so the official openai packages for Python and Node.js, and other OpenAI-compatible tools, work once you set the base URL to https://api.m-lab.app/v1 and use a Milly Lab key.

Does the API support tool calling and image input?

Not yet. The tools and tool_choice fields are accepted so SDK defaults do not fail, but they are not executed. Chat messages take text only; image parts are ignored.

Can I generate video or 3D through the API?

No. The API covers chat completions, embeddings and image generation. Video and 3D generation are not available through /v1.

Can I call a model my plan does not include?

Yes. On the Max plan and higher you can call any model that GET /v1/models lists. A model without a monthly allowance on your plan, such as Grok 4.6, is paid from your top-up balance at 2× the listed price.

Is there a free API tier?

No. Creating API keys requires the Max plan or higher. The Free plan gives 10 chat messages per model every 4 hours in the web app, without API access.

How do I pay for API usage?

API usage is paid from your plan's monthly allowance first and then from your top-up balance; prices are in US dollars. Payments are processed by SmartPay, a payment system from Tajikistan, and can be made in any currency.

Where is the full API reference?

The API guide covers streaming, embeddings, images, the usage endpoint and best practices in detail.