Gemini 3.6 Flash for Vision: Evaluation and Benchmarks

Gemini 3.6 Flash for Vision: Evaluation and Benchmarks 图片 1
Gemini 3.6 Flash for Vision: Evaluation and Benchmarks 图片 2
Gemini 3.6 Flash for Vision: Evaluation and Benchmarks 图片 3
Gemini 3.6 Flash for Vision: Evaluation and Benchmarks 图片 4

Google released Gemini 3.6 Flash on July 21, 2026, together with Gemini 3.5 Flash-Lite. Google calls 3.6 Flash its " workhorse model. " It is faster and cheaper than Gemini 3.5 Flash, and it uses about 17% fewer output tokens for the same task.

We ran both new models through our Roboflow Vision Evals to see how they do on real images and videos, not just text benchmarks. These evals are private. The majority of test images and answers are held back and never published, so no model can be tuned (benchmaxxed) on them.

TL;DR: Gemini 3.6 Flash is a good deal. It matches or leads Gemini 3.5 Flash on most image tasks. It is the best model we have tested on video, and it costs less to run. The one real step back is object detection, where it gets lazy and drops far below Gemini 3.5 Flash.

Gemini 3.6 Flash for Object DetectionObject detection is the weak spot. On mAP@50 (a standard detection score, higher is better), Gemini 3.6 Flash falls to the bottom half of the pack, well behind Gemini 3.5 Flash and even behind the cheaper Flash-Lite.

It seems as though 3.6 Flash got lazy. On many images it returns one large, loose box instead of several tight ones. It also writes malformed JSON often enough that some responses fail to parse, so you lose results even when the model clearly saw the objects.

Other users have reported the same single-box behavior since launch. The examples below show this happening: in each one it draws a handful of boxes for a scene that holds dozens.

Gemini 3.6 Flash for Vision

Outside detection, Gemini 3.6 Flash sits at or near the front.

Object detection and counting.

Counting is a surprise: Gemini 3.6 Flash leads it, just ahead of Gemini 3.5 Flash, even though it sits near the bottom on detection. So the model did not get worse at seeing objects. It just got lazy about drawing a box around each one.

Data extraction and reasoning.

Gemini 3.6 Flash ties for the top on data extraction and lands mid-pack on reasoning. The new Gemini 3.5 Flas…

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论