Skip to main content

Models

Novin Cloud provides access to dozens of language models from various providers. All of them are available through a single API with a single key.

The live model list is available in the AI → Models section of the user console.

Viewing models in the console

To see the model list in the console, you must first have an active API key and select it from the menu at the top of the page. The model list is shown based on that key's access.

Model ID​

Every model has a Model ID used in requests. IDs follow the pattern provider/model-name:

{
"model": "anthropic/claude-sonnet-4-5",
"messages": [...]
}
warning

The model ID must match the list exactly. gpt-4o without the openai/ prefix is not valid and returns a 404 error.

To get the exact list of IDs programmatically:

curl https://iapi.novin.cloud/v1/models \
-H "Authorization: Bearer $NOVIN_API_KEY"

Model specifications​

Each model's card in the console shows the following information:

Context window​

The maximum number of tokens the model can process in a single request — including the total of the input message, conversation history, and generated response.

If the total exceeds this limit, the request fails with an error. For long conversations or processing large documents, choose a model with a larger context window.

Window sizeApproximate use case
64K tokensEveryday conversations, short text
128K tokensLong documents, large codebases
200K tokensMultiple documents at once, full projects
1M tokensBooks, entire code repositories
What is a token?

A token is the unit of text processing used by models. Roughly, one token is equivalent to three to four characters of English text. Persian text usually consumes more tokens than an equivalent amount of English text.

Capabilities​

Each model may support the following capabilities:

CapabilityDescription
VisionThe ability to receive and analyze images alongside text
Function CallingThe ability to call tools and functions you define — the foundation for building agents
StreamingSending the response incrementally, word by word

Input and output pricing​

Each model has two separate prices:

  • Input price — for the tokens you send to the model (message and history)
  • Output price — for the tokens the model generates

The output price is usually three to five times the input price.

Up-to-date pricing

Prices may change. Always check the current price for each model on the Models page in the console. The figures on this page are provided only for relative comparison between models.

Providers and model families​

Anthropic — the Claude family​

Claude models excel at reasoning, long-form writing, and coding, and have a 200K-token context window.

ModelNotes
anthropic/claude-opus-4The most capable model in the family, for complex tasks
anthropic/claude-sonnet-4-5A good balance of quality and cost
anthropic/claude-haiku-4-5Fast and inexpensive, for simple, high-volume tasks
anthropic/claude-3.5-haikuThe previous generation of the fast model

OpenAI — the GPT and o-series families​

ModelNotes
openai/gpt-4oA multimodal model with image support
openai/gpt-4o-miniA lightweight, inexpensive version of gpt-4o
openai/gpt-4-turboThe previous generation, still capable
openai/o1, openai/o3Reasoning models for complex math and logic problems
openai/o4-miniA lighter reasoning model
Reasoning models

The o-series models work through reasoning steps before responding. They're more accurate for complex math, logic, and programming problems, but slower and more expensive. Don't use them for simple questions.

Google — the Gemini family​

The defining feature of this family is its very large context window (over one million tokens) — suited to processing large documents and code repositories.

ModelNotes
google/gemini-2.5-proThe most capable model in the family
google/gemini-2.5-flashFast, with the same large context window
google/gemini-2.5-flash-liteThe lightest and cheapest option

Meta — the Llama family​

Open-source models with good quality and reasonable pricing.

ModelNotes
meta-llama/llama-3.3-70b-instructGood quality at a balanced cost
meta-llama/llama-3.1-405b-instructThe largest model in the family
meta-llama/llama-3.2-11b-vision-instructSupports images

DeepSeek​

Cost-effective models focused on reasoning and coding.

ModelNotes
deepseek/deepseek-chat-v3.1The cheapest option for general tasks
deepseek/deepseek-r1A reasoning model
deepseek/deepseek-r1-distill-llama-70bA distilled, lighter version of the reasoning model

Mistral​

ModelNotes
mistralai/mistral-large-2411The flagship model
mistralai/codestral-2508Specialized for coding, 256K-token context window
mistralai/mistral-nemoLightweight and inexpensive

Qwen​

ModelNotes
qwen/qwen3-235b-a22bThe largest model in the family
qwen/qwen3-coderSpecialized for coding
qwen/qwen-plusCost-effective for everyday tasks

Model selection guide​

Your needRecommendation
Simple, high-volume tasks (classification, short summarization)Lightweight models: gpt-4o-mini, claude-haiku-4-5, gemini-2.5-flash-lite
General conversation with good qualitygpt-4o, claude-sonnet-4-5, gemini-2.5-flash
Long-form, high-quality writingclaude-sonnet-4-5, claude-opus-4
Codingcodestral-2508, qwen3-coder, claude-sonnet-4-5
Complex math and logic problemsopenai/o1, openai/o3, deepseek-r1
Processing very large documentsThe gemini-2.5 family (million-token window)
Image analysisModels with vision capability
Lowest costdeepseek-chat-v3.1, mistral-nemo, qwen-plus
A practical approach to choosing

Start with a mid-tier model. If the response quality is sufficient, try a cheaper model to find the lowest acceptable cost. If it's not sufficient, move up to a more capable model. Testing with the console chat is the fastest way to compare.

On the models page in the console, you can filter the list by provider and capability, or search by model name.

Next steps​