EmbeddingGemma 2 Is Apache 2.0. Here's Why That Matters.

Abstract tech illustration: EmbeddingGemma 2 Is Apache 2.0. Here's Why That Matters.

Your search bar, support inbox, or document Q&A runs on vectors someone else's API computed, and those vectors only work with that one model. If the vendor retires it, you pay to re-embed everything on their schedule. Google just released EmbeddingGemma 2 under Apache 2.0, and I think the license matters more than the benchmark table.

I haven't run my own benchmarks on this model yet, so I won't quote performance numbers as if I had. Here's what I can say from Google's documentation and from building AI automations for small teams.

The license is the story: what actually changed

EmbeddingGemma 2 ships under Apache 2.0, and the first EmbeddingGemma did not. Version 1 was released under the custom Gemma terms. Google's model card for version 2 lists Apache 2.0 with Google DeepMind as the author, and the developer docs say you can fine-tune the model and deploy it in your own projects and applications. According to MarkTechPost, this is a change from version 1.

Why does the difference between "Gemma terms" and Apache 2.0 matter? A March 2026 analysis from the law firm WCR.LEGAL listed three features of the old Gemma Terms of Use: a Prohibited Use Policy, a flow-down obligation to downstream users, and a unilateral termination right. Apache 2.0 has none of these. If you're a small business shipping a product on top of a model, "the licensor can terminate" is a sentence you want absent from your dependency chain.

Two caveats so you don't overread this:

  • Apache 2.0 protects the weights you already downloaded. It doesn't stop Google from releasing a future version under different terms. Pin what you use; don't assume the next release follows.
  • The official model card lists Apache 2.0 as the license, but it also states that deployments must adhere to the Gemma Prohibited Use Policy. So "Apache 2.0" is not the whole picture. Read the full card yourself before telling a client there are zero use restrictions.

Why a swap costs more than the API bill

Vectors from one embedding model are not compatible with vectors from another, so changing models means re-embedding your entire corpus. Technical sources are consistent on this (for example, TokenMix). The models place text in different vector spaces. Mix them and your similarity search returns nonsense.

I want to be honest about the cost side, because the "hidden bill" story is easy to overstate. Hosted embedding prices are low per token, and Google's hosted Gemini Embedding 2 is generally available as of April 2026. Check the current rates on the Google Cloud pricing page, since prices can change.

So the per-token price is rarely the problem. The problem is the lock-in:

  • The vendor deprecates the model you built on.
  • Your stored vectors are tied to the old model.
  • Switching means re-embedding all of it, at their timeline.
  • You also re-run any evaluation you did, because retrieval quality shifts with the model.

A small invoicing company with ten years of documents, an agency with thousands of client emails and call notes, a founder with every SOP they ever wrote: that's a lot of text to push back through a model. Sometimes it's a rounding error. Sometimes it's a project. You find out which one only when the vendor forces the question.

What EmbeddingGemma 2 actually is

EmbeddingGemma 2 is a 740M-parameter multimodal embedding model built on the Gemma 4 architecture, and you can load only the parts you need. From the official model card:

Configuration Parameters
Text only 270M
Text + image 440M
Text + audio 570M
Full multimodal 740M

The text backbone is 270M (130M transformer plus 140M embedder). Vision (170M) and audio (300M) encoders are optional. For a typical small-business use case, search over emails, contracts, and notes, text-only is what you want. The model card tells you to disable the unused encoders through SentenceTransformer's config_kwargs, so you don't pay for modalities you aren't using.

Other specs worth knowing, all from the model card:

  • Output: 768-dimensional vectors, truncatable to 512, 256, or 128 via truncate_dim.
  • Context window: 8,192 tokens, shared across all modalities. Google says that's four times larger than version 1.
  • Languages: supports 100+ languages, but the card warns performance is uneven across them.
  • Footprint: Google reports about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model (quantized, measured on a Pixel 11 Pro, per the Google Developers Blog). Those are Google's own measurements, not mine, but they suggest this runs on modest hardware.

Google also reports an MTEB multilingual v2 score of 61.36 versus 61.15 for version 1, and an MTEB Code score of 78.68 versus 68.76. Those are company-reported figures. Google's "best-in-class" language is likewise a company claim. That's exactly why the next section is about testing on your own data.

Running it: a thin wrapper and the config that matters

Keep your embedding step behind one function, store the model name and version with every vector, and the model becomes a swappable part instead of a foundation. Here's a minimal sketch based on the model card's quick start. Treat it as a starting point and check the card for the exact config_kwargs keys, since I'm not reproducing them from memory.

# embedder.py: the ONLY file that knows which model you use
from sentence_transformers import SentenceTransformer
import torch

MODEL_ID = "google/embeddinggemma-2"
MODEL_VERSION = "embeddinggemma-2"   # store this alongside every vector
DIM = 256                            # truncated from 768; see below

_model = SentenceTransformer(
    MODEL_ID,
    truncate_dim=DIM,
    model_kwargs={"torch_dtype": torch.bfloat16},  # NOT float16
    # config_kwargs=...  # disable vision/audio encoders for text-only;
    #                    # exact keys are in the official model card
)

def embed(texts: list[str]) -> list[list[float]]:
    return _model.encode(texts, normalize_embeddings=True).tolist()

And the storage side, so you always know which vector space a record lives in:

record = {
    "id": doc_id,
    "text": chunk,
    "vector": embed([chunk])[0],
    "embedding_model": MODEL_VERSION,
    "embedding_dim": DIM,
}

Three details from the model card that will save you debugging time:

  • Don't use float16. The card warns it can return NaN or silently degraded embeddings. Use bfloat16 or float32.
  • Use the task instruction prefixes for text. Omitting them still works but reduces precision. The card documents the exact prefixes.
  • Truncation is a real lever. Google says quality impact is minimal down to 256 dimensions. At 128 dimensions (1:6 compression), the MTEB multilingual score drops from 61.36 to 57.89, and 128d is best suited to text-only workloads. Smaller vectors mean a smaller index and faster search, so 256 is a sensible starting point to test.

The 30-minute lock-in audit

Write down three numbers before you pick any embedding model: which model and version you use, how many vectors you store, and what a full re-embed would cost. Do this today, regardless of whether you ever touch EmbeddingGemma.

  1. Model and version. Exact string, not "OpenAI embeddings."
  2. Vector count. Check your vector store.
  3. Re-embed cost. Document count x average tokens per document x your vendor's per-token price. Check the vendor's current pricing page rather than trusting a blog (including this one).

Illustrative math, with round numbers and not a real client: 100,000 documents at 1,000 tokens each is 100M tokens. Multiply that by your vendor's per-million-token rate and you have your re-embed bill. For most small-business corpora of this size, that's a small number, so relax. If your corpus is 100x bigger, or if re-embedding means re-validating a production search system, the dollar figure stops being the whole story. The engineering time is the cost.

If the number is small, stay on your hosted API. For a few hundred documents embedded once and never touched again, hosted is simply the convenient choice. The risk scales with corpus size and system lifespan.

If it's large, or you handle sensitive data (contracts, financial records, medical notes), self-hosting has a second benefit: your text never leaves your machine to be embedded. That makes a compliance conversation simpler. Only simpler, though. It doesn't replace one.

Then evaluate honestly. Embed a few hundred real documents, run your twenty most common searches, and compare results side by side with your current model. A small open model isn't automatically better, and your data is the only benchmark that counts. I'd want to test it on a real client corpus before promising anything.

The opinion part

Generation models change fast, and you should rent them freely. Swap Claude for something else next quarter and nothing breaks. Embeddings are different: they're a storage format, not just compute. Choosing an embedding model is closer to choosing a database than choosing a chatbot, and you wouldn't put your company's records in a format the vendor can discontinue and then charge you to migrate.

That doesn't make hosted APIs wrong. It makes defaulting to a closed, hosted-only embedding model for a long-lived system a decision you should make deliberately, not by accident because it was the fastest way to get a demo working. Apache 2.0 weights give you the option to pin the exact model and run it wherever you want. One more caution: nothing I found documents whether EmbeddingGemma 2 vectors are compatible with Gemini Embedding 2 or any other vendor's vectors. Assume they aren't, and plan any migration as a full re-embed.

Why bizflowai.io helps with this

At bizflowai.io, I build AI automations for small teams. When a project involves retrieval, I recommend a simple pattern from the first deployment: a single embedding wrapper, model name and version stored with every vector, and a re-embed script ready before anyone needs it. That way, moving to a new embedding model, open or hosted, is a planned afternoon with a side-by-side test on your own documents, not an emergency migration.


Want more like this?

I publish practical AI automation, GenAI engineering, and faceless content workflows on YouTube every week.

Subscribe to bizflowai.io on YouTube — never miss a new tutorial.

Planning an AI automation project or need a second opinion on your architecture?

Connect with me on LinkedIn — Lazar Milićević, senior engineer building practical AI automation for solopreneurs and small teams.

Visit bizflowai.io for our services, case studies, and AI consulting.

Frequently asked questions

What is EmbeddingGemma 2?

EmbeddingGemma 2 is the second generation of Google's open embedding model family. It is released under the Apache 2.0 license, so you can download the weights, run them on your own hardware, modify them, and use them commercially without asking permission. An embedding model turns text into a list of numbers called a vector, which powers search, recommendations, and document Q&A.

Why does an embedding model's license matter more than its benchmark score?

Vectors from one embedding model are not compatible with vectors from another, so switching models means re-embedding everything you've stored. With a closed, hosted-only model, the vendor can deprecate it and force you to pay that cost on their schedule. An Apache 2.0 model lets you pin the exact weights, so the same model produces the same vectors indefinitely and nobody can retire it.

How do I calculate my embedding lock-in cost?

Open your embeddings system and record three facts: the exact embedding model and version in use, how many vectors you have stored, and what it would cost to re-embed them all tomorrow. To estimate that cost, multiply your document count by your average token length by your vendor's price per token. If the result is about twelve dollars, the risk is small. If it reaches the hundreds or thousands, you have real lock-in risk.

When should I use a hosted embedding API vs. an open-weights model?

A hosted API is fine if you embed a few hundred documents once and rarely touch them again, since it needs no hardware or ops and switching risk is small. An open-weights model fits better as your vector store grows and your system lives longer, because lock-in risk scales with volume and time. Open models also let you run locally, cut per-embedding cost, and keep sensitive text on your machine.

What are the benefits of running an open-weights embedding model on your own hardware?

You can pin exact model weights so your stored vectors stay valid forever, and nobody can retire the model out from under you. Your cost per embedding drops to the electricity and hardware you already own. If you handle sensitive data like client contracts, financial records, or medical notes, your text never leaves your machine to be embedded, which simplifies compliance conversations.