Amux

Generate image

Last updated October 8, 2026

Generate images with google/gemini-nano-banana-2.1 on Gemini's native endpoint, usable with the official SDKs.

Image generation with google/gemini-nano-banana-2.1. It uses Gemini's native generateContent at the same address as a chat call, so the official SDKs work once their base URL points here.

It succeeds Nano Banana 2, with better visual quality, prompt adherence, multi-turn character consistency and text rendering. Its resolutions start at 1K — there is no 512 tier.

POSThttps://gateway.amux.ai/v1beta/models/google/gemini-nano-banana-2.1:generateContent

Authorization

header
AuthorizationstringRequired

Request

application/json
contentsarray<object>Required

Prompt and reference inputs. Text in parts[].text; references (images, video, PDF) inline in parts[].inline_data or by link in parts[].file_data. Any reference makes it image-to-image; the endpoint does not change.

›contentsarray<object>
role"user" | "model"

Who is speaking. user is you, model is the model's own earlier reply. Optional for a single-turn request.

partsarray<object>

The pieces of this turn, in the order given. Text and reference images may be mixed in one turn.

›partsarray<object>
textstring

A text piece. Several text pieces in one turn are concatenated in order.

inline_dataobject

An inline reference: an image, video or PDF. Up to 100 MB per request in total and 50 MB per PDF; use file_data for anything larger.

›inline_dataobject
mime_typestring

Media type of the input, e.g. image/png, image/jpeg, video/mp4, application/pdf.

datastring

The file bytes as base64, without the data:image/png;base64, prefix. With the prefix the upstream reads it as an invalid file.

file_dataobject

A reference supplied by link, fetched by the upstream. Suited to large files.

›file_dataobject
mime_typestring

Media type of the input, written as for inline_data.mime_type.

file_uristring

A Files API URI, or a publicly reachable HTTPS link (up to 15 MB per file). If it cannot be fetched, the upstream reports an error.

generationConfigobjectRequired

Generation settings. The fields relevant to image output are responseModalities and the size controls. The size controls accept two equivalent forms, imageConfig and responseFormat.image. Google's documentation has used each of them in different versions; Amux supports both. When both are supplied, imageConfig takes precedence, field by field.

›generationConfigobject
responseModalitiesarray<"TEXT" | "IMAGE">Required

Required: ["TEXT", "IMAGE"]. Without it this endpoint returns text only.

responseFormatobject

Output format. Image settings live under image.

›responseFormatobject
imageobject

Size controls for the image. Both may be omitted, in which case the upstream picks its default.

›imageobject
aspectRatio"1:1" | "1:4" | "1:8" | "2:3" | "3:2" | "3:4" | "4:1" | "4:3" | "4:5" | "5:4" | "8:1" | "9:16" | "16:9" | "21:9"

Aspect ratio. This model accepts the 14 values below. An out-of-range value is dropped and noted in amux.notes, and the model falls back to its default; this endpoint, /v1/images/* and /v1/tasks all behave the same.

imageSize"1K" | "2K" | "4K"

Resolution tier. Out-of-range values are handled the same way as aspectRatio.

toolsarray<object>

Grounding with Google Search: [{ "google_search": {} }]. Billed per search query; the actual list is in groundingMetadata.webSearchQueries.

Response

200response

The result. Images and text are interleaved in parts.

Turning image output on

generationConfig.responseModalities must be ["TEXT", "IMAGE"]. Without it the same endpoint returns text only.

What this model accepts

Aspect ratios (14):

1:1 · 2:3 / 3:2 · 3:4 / 4:3 · 4:5 / 5:4 · 9:16 / 16:9 · 21:9 · 1:4 / 4:1 · 1:8 / 8:1

Resolutions: 1K · 2K · 4K, defaulting to 1K when omitted

A value outside those sets is dropped and noted in amux.notes, after which the model falls back to its own default. This native endpoint shares one image pipeline with the OpenAI-compatible image endpoint and Amux Tasks, so all three behave the same.

An OpenAI-style pixel string in size (for example 1024x1024) is folded into an aspect ratio and a tier, constrained to the sets above, and a lossy fold is noted in amux.notes. Gemini's native body has no such field — on a native call, write the aspect ratio and the tier directly.

Reference inputs

Image-to-image and text-to-image share this endpoint: put the references into contents[].parts. Besides images, this model also accepts video and PDF — for example, drawing in the style of a clip, or following a product brief in a PDF.

There are two ways to supply them:

  • Inline: inline_data, where data is raw base64 without the data: prefix and mime_type is the actual type (such as image/png, video/mp4 or application/pdf). Up to 100 MB per request in total, and up to 50 MB for any single PDF; anything larger is rejected.
  • By link: file_data, where file_uri is a Files API URI or a publicly reachable HTTPS link and mime_type is as above. A public link may be up to 15 MB per file; the upstream fetches it and reports an error if it cannot. Prefer this for large files.

Up to 14 reference images, of which at most 10 may be object references and 4 character references. Types other than image, video and PDF are forwarded as is, with a note in amux.notes.

Thinking

The model thinks before it draws. The levels are minimal, medium and high, with medium as the default. On the native endpoint set it in generationConfig.thinkingConfig.thinkingLevel; it is passed to the upstream as is. Deeper thinking gives steadier composition and text, at the cost of latency and thinking tokens.

What drives the cost

How output is billed depends on the provider: for the same model, some providers charge per token and others per image. The provider's price list on the model page is authoritative. All three call paths — this endpoint, /v1/images/* and /v1/tasks — share one image pipeline, so a given provider bills the same on all of them.

When billing is per token, image tokens and text tokens are priced differently. The two are reported separately in usageMetadata.candidatesTokensDetails, and Amux bills them as separate line items rather than charging image tokens at the text rate. Thinking tokens are billed as text output, and video and PDF tokens as text input.

How many tokens one image costs depends on the resolution: 1120 tokens at 1K, 1680 at 2K, 3780 at 4K.

For the rates themselves see the model page and pay-as-you-go pricing. They are deliberately not repeated here — providers and discounts change, and a copy in the docs would drift away from the invoice.

tools: [{ "google_search": {} }] lets the model search before generating — useful when the picture has to reflect current facts.

The camelCase googleSearch that the SDKs emit is accepted too. Through the OpenAI-compatible image endpoint or Amux Tasks, the equivalent is web_search: true.

Billing is per search query, and one call may issue several. The count is the de-duplicated candidates[].groundingMetadata.webSearchQueries; empty queries are not charged.

Streaming is not supported

Gemini streams by switching the action (:streamGenerateContent), and the official documentation gives no syntax for streaming images, so streaming requests to this model are rejected. Retry without streaming.

One call returns one image; the request body has no batch field.

Other ways to call this model

The model is also reachable through the OpenAI-compatible image endpoint and Amux Tasks. The ranges listed on this page apply on those converted paths too.

Errors

Errors use Gemini's error body. See Errors and retries for status values and retry semantics.

cURL
curl https://gateway.amux.ai/v1beta/models/google/gemini-nano-banana-2.1:generateContent \
  -H "x-goog-api-key: $AMUX_API_KEY" \
  -d '{
    "contents": [
      {
        "role": "user",
        "parts": [
          {
            "text": "<string>",
            "inline_data": {
              "mime_type": "image/png",
              "data": "<string>"
            },
            "file_data": {
              "mime_type": "video/mp4",
              "file_uri": "https://example.com/clip.mp4"
            }
          }
        ]
      }
    ],
    "generationConfig": {
      "responseModalities": [
        "TEXT",
        "IMAGE"
      ],
      "responseFormat": {
        "image": {
          "aspectRatio": "1:1",
          "imageSize": "1K"
        }
      }
    },
    "tools": [
      {}
    ]
  }'
{
  "candidates": [
    {
      "content": {
        "parts": [
          {
            "text": "<string>",
            "inlineData": {
              "mimeType": "image/png",
              "data": "<string>"
            }
          }
        ]
      }
    }
  ],
  "usageMetadata": {
    "promptTokenCount": 0,
    "candidatesTokenCount": 0,
    "totalTokenCount": 0,
    "promptTokensDetails": [
      {
        "modality": "TEXT",
        "tokenCount": 0
      }
    ],
    "candidatesTokensDetails": [
      {
        "modality": "IMAGE",
        "tokenCount": 0
      }
    ]
  }
}