Google 于 2026 年 10 月 6 日发布 Nano Banana 2.1,现在已可通过 Gemini API 以 gemini-nano-banana-2.1 调用。它是 Nano Banana 2(Gemini 3.1 Flash Image)的更新版,带来了更好的画面质量、mask 式编辑,以及更强的跨轮角色一致性。在两者共有的所有分辨率下,它的单张图片成本也只有 Nano Banana 2 的一半。
本文从空终端开始,一步步带你生成并保存第一张图片,随后覆盖宽高比、2K/4K 输出、图像编辑、多轮修改、参考图、搜索接地和费用。下文每个请求都可以在 Apifox 中保存并反复重放,方便你把 2.1 与当前在用的模型做对比。
想先了解背景?可以看《Nano Banana 2.1 是什么》,里面讲了有哪些变化,以及 Google 尚未公布的内容。
准备工作
| 项目 | 值 |
|---|---|
| 前置 URL | https://generativelanguage.googleapis.com/v1beta |
| 鉴权 header | x-goog-api-key: $GEMINI_API_KEY |
| 模型 ID | gemini-nano-banana-2.1 |
| 接口 | POST /v1beta/interactions |
| 分辨率 | 1K(默认)、2K、4K |
| 输入 | 文本、图片(最多 14 张参考图)、视频 |
| Python SDK | pip install google-genai |
| JavaScript SDK | npm install @google/genai |
所有示例都使用 Interactions API,Google 的图像生成文档中针对 2.1 用的也是它。
第 1 步:获取 Gemini API key
- 打开 Google AI Studio 并登录。
- 进入 Get API key,在一个 Google Cloud 项目下创建 key。
- 为该项目的开启结算。Google 定价页把 Nano Banana 2.1 的免费档标为“Not available”,因此调用 API 必须使用已开启结算的项目。
- 导出 key:
export GEMINI_API_KEY="your-key-here"
Google 把 2.1 链到了 AI Studio 试玩页,用它测试 prompt 很方便,但 Google 没有公布免费账号在那里能生成多少。我们在《如何免费使用 Nano Banana 2.1》中整理了这条路径,《如何获取 Gemini API key》则更详细地讲解了 key 的配置流程。
第 2 步:生成第一张图片
curl
curl -s -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-nano-banana-2.1",
"input": [
{"type": "text", "text": "Create a picture of a nano banana dish in a fancy restaurant with a Gemini theme"}
]
}'
响应是 JSON。图片以 base64 数据的形式放在 image 内容块中返回,所以你需要一个 SDK(或一段脚本)来解码。
Python
from google import genai
import base64
client = genai.Client() # reads GEMINI_API_KEY
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input="Create a picture of a nano banana dish in a fancy restaurant with a Gemini theme",
)
with open("generated_image.png", "wb") as f:
f.write(base64.b64decode(interaction.output_image.data))
interaction.output_image 返回最后一个生成的图片块,其 data 字段是 base64,写入文件前需要先解码。
JavaScript
import { GoogleGenAI } from "@google/genai";
import * as fs from "node:fs";
const ai = new GoogleGenAI({});
const interaction = await ai.interactions.create({
model: "gemini-nano-banana-2.1",
input: "Create a picture of a nano banana dish in a fancy restaurant with a Gemini theme",
});
const image = interaction.output_image;
if (image) {
fs.writeFileSync("nano-banana.png", Buffer.from(image.data, "base64"));
}
Gemini 3 图像模型会先思考再作画。模型在规划构图时最多会产生两张中间的“思考图像”,这部分不计费,而且在 API 中无法关闭 thinking。
第 3 步:设置宽高比、分辨率和纯图片输出
用 response_format 控制输出。设置 "type": "image" 后只返回图片、不带对话文本,响应体更小,解析也更简单。
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input="A product shot of a matte black coffee grinder on a marble counter",
response_format={
"type": "image",
"mime_type": "image/png",
"aspect_ratio": "16:9",
"image_size": "2K",
},
)
几点需要了解:
image_size接受1K、2K和4K。请使用大写的K;像2k这样的小写值会被拒绝。- 2.1 没有 512px 档位,那一档属于 Nano Banana 2。
- 支持的宽高比包括
1:1、2:3、3:2、3:4、4:3、4:5、5:4、9:16、16:9、21:9,以及更极端的1:4、4:1、1:8和8:1。2K 下的 16:9 图片为 2752x1536,4K 下为 5504x3072。 - 如果既要文本又要图片,传入列表即可:
response_format=[{"type": "text"}, {"type": "image"}]。
第 4 步:编辑已有图片
把图片以 base64 的 image 块形式,与文本指令一起发送:
with open("living_room.png", "rb") as f:
image_b64 = base64.b64encode(f.read()).decode("utf-8")
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input=[
{"type": "text", "text": "Using the provided image of a living room, change only the blue sofa to be a vintage, brown leather chesterfield sofa. Keep the rest of the room, including the pillows on the sofa and the lighting, unchanged."},
{"type": "image", "data": image_b64, "mime_type": "image/png"},
],
)
无需 mask 文件的 mask 式局部重绘
这里不需要单独上传 mask,用文字描述即可定义遮罩。Google 给出的模板是:
Using the provided image, change only the [specific element] to [new element/description]. Keep everything else in the image exactly the same, preserving the original style, lighting, and composition.
只写一个元素,具体描述替换内容,并重申哪些部分必须保持不变。像“让沙发更好看”这种含糊的 prompt,会诱使模型重画整个房间。
第 5 步:用多轮编辑反复微调
多轮编辑是 Google 推荐的图片精修方式。把上一轮交互的 id 作为 previous_interaction_id 传入,并且只发送你想要的改动:
interaction_2 = client.interactions.create(
model="gemini-nano-banana-2.1",
input="Update this infographic to be in Spanish. Do not change any other elements of the image.",
previous_interaction_id=interaction.id,
response_format={"type": "image", "mime_type": "image/png", "aspect_ratio": "16:9", "image_size": "2K"},
)
走 REST 时,同一个字段放在请求 body 中:"previous_interaction_id": "<PREVIOUS_INTERACTION_ID>"。2.1 改进的多轮角色一致性正是在这里体现价值:角色或产品在多轮编辑后仍应保持可辨识。
第 6 步:组合最多 14 张参考图
把更多 image 块追加到 input 列表中。对于 2.1,Google 文档说明最多支持 10 张高保真物体图和 4 张角色图,合计 14 张:
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input=[
{"type": "text", "text": "An office group photo of these people, they are making funny faces."},
{"type": "image", "data": person_1_b64, "mime_type": "image/png"},
{"type": "image", "data": person_2_b64, "mime_type": "image/png"},
{"type": "image", "data": person_3_b64, "mime_type": "image/png"},
],
response_format={"type": "image", "aspect_ratio": "5:4", "image_size": "2K"},
)
参考图会计入输入 token,而 2.1 的输入价格是 Nano Banana 2 的三倍。参考图用得越多,价差就越小,因此切换前请先实测。
第 7 步:用 Google Search 为图片接地
对于依赖实时事实的图片,例如天气图或近期事件,可以加上搜索工具:
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input="A detailed painting of a Timareta butterfly resting on a flower",
tools=[{"type": "google_search", "search_types": ["web_search", "image_search"]}],
)
只写 {"type": "google_search"} 时使用网页搜索。加上 image_search 后,模型可以使用网络图片作为视觉上下文,这一能力目前只有 2.1 和 Nano Banana 2 支持。接地不能使用搜索获得的真人图像。如果你把接地结果展示给用户,Google 要求你展示返回的 search_suggestions,它来自 google_search_result 步骤。
第 8 步:在 Apifox 中测试
生成图片用脚本就够了。但要核查 prompt、key 和模型是否仍按预期工作,保存成请求会更快。在 Apifox 中:
- 创建一个环境,添加变量
GEMINI_API_KEY,填入你的 key。 - 新建请求:
POST https://generativelanguage.googleapis.com/v1beta/interactions。 - 添加 header
x-goog-api-key: {{GEMINI_API_KEY}}。 - 粘贴包含
model、input和response_format的 JSON body,然后点击发送。 - 添加一个后置操作脚本,包含两条断言:
pm.test("status is 200", () => {
pm.response.to.have.status(200);
});
pm.test("response contains an image block", () => {
const steps = pm.response.json().steps || [];
const hasImage = steps.some(s =>
s.type === "model_output" &&
(s.content || []).some(c => c.type === "image" && c.data)
);
pm.expect(hasImage).to.be.true;
});
- 复制该请求,把
model改成gemini-3.1-flash-image,用同一个 prompt 分别发送。
这样你就得到了 2.1 与 Nano Banana 2 的并排对比,每次运行都会展示状态码、耗时和响应体大小。更全面的功能对比,可参阅《Nano Banana 2.1 vs Nano Banana 2 vs Pro》。
费用
付费档价格,来自 Google 定价页,2026 年 10 月 7 日:
| 模型 | 输入 / 1M | 1K 图片 | 2K 图片 | 4K 图片 | 批量 1K |
|---|---|---|---|---|---|
| Nano Banana 2.1 | $1.50 | $0.0336 | $0.0504 | $0.0756 | $0.0168 |
| Nano Banana 2 | $0.50 | $0.067 | $0.101 | $0.151 | $0.034 |
| Nano Banana Pro | $2.00 | $0.134 | $0.134 | $0.24 | $0.067 |
2.1 的图片输出为 $30/1M token。1K 图片为 1,120 token,2K 为 1,680,4K 为 2,520,单张价格就是由此换算而来。文本与 thinking 输出为 $7.50/1M。搜索接地每月含 5,000 次免费请求,由 Gemini 3.x 系列模型共享,超出后 $14/1000 次。
算例:用简短文本 prompt(每次约 100 输入 token)生成 1,000 张 2K 产品图。
- 输出:1,000 × $0.0504 = $50.40
- 输入:100,000 token × $1.50 / 1M = $0.15
- 合计:约 $50.55,另加 thinking 文本产生的 token
同样的任务在 Nano Banana 2 上约需 $101。走 Batch API 时,2.1 的 2K 价格降到 $0.0252,只要你能等待,输出部分就降到 $25.20。批量任务用最长 24 小时的周转时间换取更高的速率上限。Nano Banana 2 的更多定价细节,见我们的 Nano Banana 2 API 定价拆解。
常见错误
以下属于 Gemini API 的通用行为,并非 2.1 专属的已记录错误:
| 错误 | 可能原因 | 解决办法 |
|---|---|---|
400 INVALID_ARGUMENT |
image_size 用了小写、宽高比不受支持、input 格式有误 |
使用 2K 而不是 2k;核对宽高比列表 |
403 PERMISSION_DENIED |
key 无效,或项目未开启结算 | 检查 key 并开启结算 |
404 NOT_FOUND |
模型 ID 拼写错误 | 严格使用 gemini-nano-banana-2.1 |
429 RESOURCE_EXHAUSTED |
触发速率限制 | 退避后重试,或改用 Batch |
500 / 503 |
服务端临时故障 | 使用指数退避重试 |
| 返回 200 但没有图片 | prompt 被拦截,或只返回了文本 | 设置 "type": "image" 并改写 prompt |
常见问题
Nano Banana 2.1 API 有免费档吗? 没有。Google 把 API 免费档标为“Not available”。AI Studio 试玩页虽链到 2.1,但 Google 没有公布它的免费额度。
2.1 能直接替换 Nano Banana 2 吗? 大体可以,把模型 ID 换成 gemini-nano-banana-2.1 即可。但你会失去 512px 档位,且输入 token 更贵,因此参考图密集的编辑需要先做成本核算。
生成的图片带水印吗? 带。所有输出都包含 SynthID 水印。
可以用视频生成图片吗? 可以。2.1 支持在文本 prompt 之外传入 video 输入块,例如 YouTube URL。
旧的 Nano Banana 2 API 指南还适用吗? 概念仍然适用。早期模型的用法见我们的 Nano Banana 2 API 指南。
小结
Nano Banana 2.1 承诺以 Nano Banana 2 一半的单张价格换来更好的画质,并且沿用相同的 Interactions API 结构。先拿到已开启结算的 key,生成一张图片,然后把请求连同状态码和图片断言一起保存到 Apifox 中。再用旧模型 ID 复制一份,一个下午就能判断 2.1 是否值得为你的 prompt 切换过去。
HiFox:将 Agent 变成真正的队友
另外,我们也在思考,AI 如何从个人提效走进团队协作。
HiFox 是一个让人和 AI Agent 在同一个工作现场协作的平台:你可以像给同事分派任务一样指派 Agent,在任务看板中跟踪进度、查看结果,让 Agent 成为团队里的队友。
👉 立即体验 HiFox:https://hifox.com
AI Coding 交流群
如果你也在用 AI 写代码,或者正在研究 Cursor、Claude Code 这些工具,
欢迎加入以下交流群。群里平时会聊一些 AI 编程的实际用法、开发工作流,还有各种新工具和新玩法。