Generate image
Last updated October 8, 2026
Generate images with google/gemini-nano-banana-2.1 on Gemini's native endpoint, usable with the official SDKs.
Image generation with google/gemini-nano-banana-2.1. It uses Gemini's native generateContent at the
same address as a chat call, so the official SDKs work once their base URL points here.
It succeeds Nano Banana 2, with better visual quality, prompt adherence, multi-turn character consistency
and text rendering. Its resolutions start at 1K — there is no 512 tier.
https://gateway.amux.ai/v1beta/models/google/gemini-nano-banana-2.1:generateContentAuthorization
headerAuthorizationstringRequiredRequest
application/jsoncontentsarray<object>RequiredPrompt and reference inputs. Text in parts[].text; references (images, video, PDF) inline in parts[].inline_data or by link in parts[].file_data. Any reference makes it image-to-image; the endpoint does not change.
›contentsarray<object>
role"user" | "model"Who is speaking. user is you, model is the model's own earlier reply. Optional for a single-turn request.
partsarray<object>The pieces of this turn, in the order given. Text and reference images may be mixed in one turn.
›partsarray<object>
textstringA text piece. Several text pieces in one turn are concatenated in order.
inline_dataobjectAn inline reference: an image, video or PDF. Up to 100 MB per request in total and 50 MB per PDF; use file_data for anything larger.
›inline_dataobject
mime_typestringMedia type of the input, e.g. image/png, image/jpeg, video/mp4, application/pdf.
datastringThe file bytes as base64, without the data:image/png;base64, prefix. With the prefix the upstream reads it as an invalid file.
file_dataobjectA reference supplied by link, fetched by the upstream. Suited to large files.
›file_dataobject
mime_typestringMedia type of the input, written as for inline_data.mime_type.
file_uristringA Files API URI, or a publicly reachable HTTPS link (up to 15 MB per file). If it cannot be fetched, the upstream reports an error.
generationConfigobjectRequiredGeneration settings. The fields relevant to image output are responseModalities and the size controls. The size controls accept two equivalent forms, imageConfig and responseFormat.image. Google's documentation has used each of them in different versions; Amux supports both. When both are supplied, imageConfig takes precedence, field by field.
›generationConfigobject
responseModalitiesarray<"TEXT" | "IMAGE">RequiredRequired: ["TEXT", "IMAGE"]. Without it this endpoint returns text only.
responseFormatobjectOutput format. Image settings live under image.
›responseFormatobject
imageobjectSize controls for the image. Both may be omitted, in which case the upstream picks its default.
›imageobject
aspectRatio"1:1" | "1:4" | "1:8" | "2:3" | "3:2" | "3:4" | "4:1" | "4:3" | "4:5" | "5:4" | "8:1" | "9:16" | "16:9" | "21:9"Aspect ratio. This model accepts the 14 values below. An out-of-range value is dropped and noted in amux.notes, and the model falls back to its default; this endpoint, /v1/images/* and /v1/tasks all behave the same.
imageSize"1K" | "2K" | "4K"Resolution tier. Out-of-range values are handled the same way as aspectRatio.
toolsarray<object>Grounding with Google Search: [{ "google_search": {} }]. Billed per search query; the actual list is in groundingMetadata.webSearchQueries.
Response
200responseThe result. Images and text are interleaved in parts.
Turning image output on
generationConfig.responseModalities must be ["TEXT", "IMAGE"]. Without it the same endpoint
returns text only.
What this model accepts
Aspect ratios (14):
1:1 · 2:3 / 3:2 · 3:4 / 4:3 · 4:5 / 5:4 · 9:16 / 16:9 · 21:9 · 1:4 / 4:1 · 1:8 / 8:1
Resolutions: 1K · 2K · 4K, defaulting to 1K when omitted
A value outside those sets is dropped and noted in amux.notes, after which the model falls back
to its own default. This native endpoint shares one image pipeline with the
OpenAI-compatible image endpoint and
Amux Tasks, so all three behave the same.
An OpenAI-style pixel string in size (for example 1024x1024) is folded into an aspect ratio and a
tier, constrained to the sets above, and a lossy fold is noted in amux.notes. Gemini's native body
has no such field — on a native call, write the aspect ratio and the tier directly.
Reference inputs
Image-to-image and text-to-image share this endpoint: put the references into contents[].parts.
Besides images, this model also accepts video and PDF — for example, drawing in the style of a clip,
or following a product brief in a PDF.
There are two ways to supply them:
- Inline:
inline_data, wheredatais raw base64 without thedata:prefix andmime_typeis the actual type (such asimage/png,video/mp4orapplication/pdf). Up to 100 MB per request in total, and up to 50 MB for any single PDF; anything larger is rejected. - By link:
file_data, wherefile_uriis a Files API URI or a publicly reachable HTTPS link andmime_typeis as above. A public link may be up to 15 MB per file; the upstream fetches it and reports an error if it cannot. Prefer this for large files.
Up to 14 reference images, of which at most 10 may be object references and 4 character references.
Types other than image, video and PDF are forwarded as is, with a note in amux.notes.
Thinking
The model thinks before it draws. The levels are minimal, medium and high, with medium as the default.
On the native endpoint set it in generationConfig.thinkingConfig.thinkingLevel; it is passed to the upstream
as is. Deeper thinking gives steadier composition and text, at the cost of latency and thinking tokens.
What drives the cost
How output is billed depends on the provider: for the same model, some providers charge per token
and others per image. The provider's price list on the model page is authoritative. All three call
paths — this endpoint, /v1/images/* and /v1/tasks — share one image pipeline, so a given provider
bills the same on all of them.
When billing is per token, image tokens and text tokens are priced differently. The two are
reported separately in usageMetadata.candidatesTokensDetails, and Amux bills them as separate line
items rather than charging image tokens at the text rate. Thinking tokens are billed as text output, and
video and PDF tokens as text input.
How many tokens one image costs depends on the resolution: 1120 tokens at 1K, 1680 at 2K, 3780 at 4K.
For the rates themselves see the model page and pay-as-you-go pricing. They are deliberately not repeated here — providers and discounts change, and a copy in the docs would drift away from the invoice.
Grounding with Google Search
tools: [{ "google_search": {} }] lets the model search before generating — useful when the picture
has to reflect current facts.
The camelCase googleSearch that the SDKs emit is accepted too. Through the
OpenAI-compatible image endpoint or Amux Tasks, the equivalent
is web_search: true.
Billing is per search query, and one call may issue several. The count is the de-duplicated
candidates[].groundingMetadata.webSearchQueries; empty queries are not charged.
Streaming is not supported
Gemini streams by switching the action (:streamGenerateContent), and the official documentation gives
no syntax for streaming images, so streaming requests to this model are rejected. Retry without streaming.
One call returns one image; the request body has no batch field.
Other ways to call this model
The model is also reachable through the OpenAI-compatible image endpoint and Amux Tasks. The ranges listed on this page apply on those converted paths too.
Errors
Errors use Gemini's error body. See Errors and retries for status values and
retry semantics.
curl https://gateway.amux.ai/v1beta/models/google/gemini-nano-banana-2.1:generateContent \
-H "x-goog-api-key: $AMUX_API_KEY" \
-d '{
"contents": [
{
"role": "user",
"parts": [
{
"text": "<string>",
"inline_data": {
"mime_type": "image/png",
"data": "<string>"
},
"file_data": {
"mime_type": "video/mp4",
"file_uri": "https://example.com/clip.mp4"
}
}
]
}
],
"generationConfig": {
"responseModalities": [
"TEXT",
"IMAGE"
],
"responseFormat": {
"image": {
"aspectRatio": "1:1",
"imageSize": "1K"
}
}
},
"tools": [
{}
]
}'{
"candidates": [
{
"content": {
"parts": [
{
"text": "<string>",
"inlineData": {
"mimeType": "image/png",
"data": "<string>"
}
}
]
}
}
],
"usageMetadata": {
"promptTokenCount": 0,
"candidatesTokenCount": 0,
"totalTokenCount": 0,
"promptTokensDetails": [
{
"modality": "TEXT",
"tokenCount": 0
}
],
"candidatesTokensDetails": [
{
"modality": "IMAGE",
"tokenCount": 0
}
]
}
}