Article

Jev and Gemini 3.8 Flash: Matching SQL, Slower NL-to-SQL

Jev still owns the NL-to-SQL judgments. Gemini 3.8 Flash matched GPT-6 SQL on the same ClickHouse demo run, and mean SQL generation rose from 5.1s to 10.4s.

QueryPanel Team
6 min read
nl-to-sqltext-to-sqljevtypesafegeminilatencyengineering

Jev already does the yes/no work in our NL-to-SQL pipeline. We pointed the writing at Gemini 3.8 Flash, left GPT-6 as the default, and ran the same ClickHouse demo. The SQL matched. The writing got slower.


Short answer: On 1 October 2026 we asked the same order questions twice, with Jev on both times. GPT-6 wrote the SQL in one run. Gemini 3.8 Flash wrote it in the other. Both runs returned the right numbers. Writing the SQL averaged 5.1s on GPT-6 and 10.4s on Gemini. A full answer averaged 11.7s vs 18.1s. GPT-6 stays the production default.

Key takeaways

  • Jev's judgment steps did not move. Intent stayed near 300ms and schema linking near 280ms on both runs.
  • Gemini 3.8 Flash wrote SQL that passed the same seeded checks as GPT-6, including tenant filters and date cutoffs.
  • The extra time is in writing SQL. The demo took 298s with Gemini and 242s with GPT-6.
  • We left Gemini's thinking level at the API default, medium. We did not retry at low.
  • Gemini 4 Argon is announced and not on the public Gemini API. This post has no Argon numbers.

The paired run

Same questions, same orders, same tenant. Jev made the judgments both times. One run used the production GPT-6 models. The other used Gemini 3.8 Flash to write SQL, charts, and the LLM fallback. We did not change the default.

GPT-6Gemini 3.8 Flash
SQL generation, mean5.1s10.4s
Full answer, mean11.7s18.1s
Whole demo242s298s
Right numbersyesyes

Reflection stayed on the LLM path, about 4.6s with GPT-6 and 4.0s with Gemini. The earlier Jev and GPT-6 post caught a fast skip near 300ms. This demo did not. Both writers paid for a full check of the SQL.

What Jev still did

Guardrail, intent, schema pruning, and follow-up routing stay on Jev, TypeSafe's System One model. The wiring is the same as in the GPT-6 post: Jev picks from options we define, our code maps the probabilities, and a low-confidence or timed-out call falls back to the LLM.

On this demo that fallback writer is the thing we swapped. When the provider is OpenAI, the fallback is GPT-6. When we set the provider to Google, the fallback is Gemini 3.8 Flash. Jev does not write the SQL.

Where Gemini 3.8 Flash fits

Gemini 3.8 Flash (gemini-3.8-flash) is generally available. Google's docs, checked 1 October 2026, say the default thinking level is medium, and minimal is rejected. We did not send a thinking override, so this bench is the default.

For the comparison only, these roles all used that id:

RoleProduction defaultThis Gemini run
SQL generationgpt-6-lunagemini-3.8-flash
Chart generationgpt-6-lunagemini-3.8-flash
Guardrail LLM fallbackgpt-6-lunagemini-3.8-flash
Query rewritergpt-6-lunagemini-3.8-flash
Schema linker LLM fallbackgpt-6.1-solgemini-3.8-flash

The SQL

Both models returned the right figures, kept the tenant filter, and used a half-open date range instead of a closed BETWEEN that drops the last day. The wording differed. Gemini leaned on ClickHouse count() and avg(). GPT-6 leaned on COUNT(*) and AVG(). The numbers matched anyway.

One question took 22s to write on Gemini and 5s on GPT-6, for the same result. That was a single sample. Set it aside and Gemini is still slower: 9.1s vs 5.1s.

Why GPT-6 stays the default

The demo did not give a reason to switch. Accuracy on this seed was a tie. Latency favored GPT-6, mostly inside SQL generation. Customers on the React SDK and callers of qp.ask() / POST /v2/query keep the same request and response. The model choice is server-side. Tenant isolation is still the prompt, the reflection pass, and the check before execution. Jev can only shorten the path when it is already confident.

Headless stays the same boundary: the Node SDK runs the SQL in your process. The model never receives query results or database credentials.

Gemini 4 Argon

Google announced Gemini 4 Argon on 30 September 2026. The public note says it is rolling out to cyber defenders in the Fairwind program first, and that a public API id for developers is later. We do not have that id, so we did not call it.

When an id is on the Gemini API, the repeat is the same demo: Jev on, every generative role set to that id, GPT-6 left as the default unless the numbers say otherwise.

FAQ

Do embed customers need a new SDK?

No. The React SDK, the Node SDK, and POST /v2/query keep the same shapes. This comparison did not change the default model those paths use.

Did Gemini see row data?

No. Gemini received the same generation context the SQL writer already gets: the question, schema chunks, and the plan. Execution stays in the SDK callback against your database. Jev's context is unchanged from the previous post.

Why was one Gemini answer so slow?

We don't know from a single sample. Thinking was at Google's default medium, and the writer drafts a few SQL candidates before keeping one. A follow-up would set thinking to low and repeat the questions. We did not do that here.

What happens if the Google call fails?

The request fails with that provider error. It does not switch itself over to GPT-6.

When do we bench Argon?

When Google publishes a model id on the Gemini API. Until then, Argon is a name in a launch post, not a number in this comparison.


Further reading