Google Will Pay for Your Private Code. Read the Terms First.

Developer reviewing source code on a laptop screen while deciding whether to sell a private repository

Google is expanding a pilot that pays developers and small businesses for non-public code and other data. If you run a small shop with a private repo that took two years to build, a "get paid for it" offer is tempting, and it raises questions nobody has published answers to yet: what exactly are you selling, what rights are you granting, and what else of yours is already leaking through the AI tools you use every day?

This post separates what's confirmed from what's inference, gives you a pre-submission checklist, and shows how to audit your own AI workflows so you decide what leaves your systems instead of finding out later.

What Google has actually confirmed

Google told VentureBeat it is expanding a pilot, called the "Content Offer Pilot", that pays developers and small businesses for specialized, non-public content, including proprietary code repositories, to improve its products and services. Prospective partners submit selected material through an online portal. Participants choose what to offer and propose a price, and according to Google a deal goes ahead only if both sides agree.

Other details from VentureBeat's report:

  • "Hundreds" of third-party developers and small businesses in roughly 100 countries have participated so far.
  • More than 90% of participating partners said they're interested in taking part again.
  • The expansion builds on an earlier effort. In June, Google said it had made deals for specialized, non-public material, including creative and educational content.

Google's own pages back up the framing. Its AI partnerships page says it trains primarily on publicly available, crawlable web data and has been piloting a program to pay for non-public content in a range of media formats. It lists three content areas: closed and offline datasets, enhanced metadata and signals, and real-time structured factual information for verification. It says it is prioritizing science, mathematics and learning contexts. Google's public policy blog similarly says it has entered into deals to pay for access to and delivery of specialized, non-public content.

That's the confirmed surface. Everything below it is thinner than the headline suggests.

What nobody has disclosed

According to the same reporting, Google did not provide payment ranges or say how many additional partners it expects to accept. It has also not said which products will use the code obtained through the pilot, how much it pays, or what rights it receives under a typical agreement.

Read that list as a due-diligence checklist. For a code sale, these are the terms that matter most:

Question Status
What does it pay? Not disclosed
Which products use the code? Not specified by Google
Is it used for model training? Not stated by Google; outlets have inferred it
What rights does Google get in a typical deal? Not published
Retention, opt-out, training limits? Not published
Eligibility criteria? Not published

One caution on the training question. Reports that this code feeds Gemini or coding agents are inference by outlets like Android Authority and Digital Trends, not a Google statement. Google's own framing is "improve our products and services." Don't assume either "it trains the model" or "it doesn't" until the contract tells you.

The June precedent: what the earlier email said

The June outreach gives a partial picture. 404 Media reported in early June 2026 that Google had emailed select Play Store developers an invitation to a "confidential content offer pilot" to buy access to their codebases. According to the email text reproduced by 404 Media:

  • The offer was non-exclusive.
  • Developers would keep 100% of their IP and the right to monetize their data elsewhere.
  • Google said real-world code is useful for understanding complex logic and for developing coding evals and benchmarks.
  • The email did not mention AI. A link in it led to a Google page about "partnerships to improve our AI products".

Two caveats. First, no source explicitly confirms that the expanded Content Offer Pilot is the same program as the June "confidential content offer pilot". The names are nearly identical, but treat them as possibly related, not confirmed identical. Second, the "100% IP, non-exclusive" language is from that June email, not from a published contract template for the expanded program.

For the expanded pilot, the public evidence on ownership is one anonymous testimonial supplied by Google, in which a startup CEO said the startup could keep ownership of its code. That's a vendor-selected quote, not a contract term. Ownership of the code and the license you grant to use it are different things. You can "keep ownership" while granting broad usage rights, and the usage rights are what determine your exposure.

Is selling code a good idea? A decision framework

Not "yes" or "no" but "for which repo, under which terms." Here's how I'd evaluate any offer, from this program or a similar one.

Step 1: Classify the repo before you think about price.

  • Never sell: code containing customer data, client-owned IP, secrets, or anything under NDA or a client contract with IP assignment. You may not have the right to license it at all.
  • Sell with caution: the core of your product's competitive edge. A one-time payment rarely compensates for a competitor's model getting better at your niche.
  • Candidate: internal tooling, old or abandoned projects, utility libraries, and prototypes with no competitive value and no third-party obligations.

Step 2: Check what you actually own. Open-source dependencies carry their own licenses. Contractor-written code needs a signed assignment. Code written for a client is usually theirs. Vendored third-party code is not yours to license. If you can't answer "who owns this line?" for the whole repo, you're not ready to sell it.

Step 3: Get the missing terms in writing. At minimum, ask for answers to:

  1. Permitted uses: training, evaluation and benchmarking, product features, human review?
  2. Retention: how long is the data kept, and can you require deletion?
  3. Scope of the license: exclusive or non-exclusive, perpetual or time-limited, sublicensable?
  4. Derivative outputs: can the system reproduce your code verbatim or near-verbatim?
  5. Confidentiality: who inside the buyer can see it, and under what controls?
  6. Payment: lump sum, recurring, or tied to use?

Since you propose the price, you also need a floor. Google hasn't published payment ranges, so there's no public benchmark for code. Anchor on your own numbers: what the repo cost to build, and what a competitor would pay for it. The Reddit-style forum-data deals aren't a valid comparison; that's a different asset type with unverified figures.

Step 4: Talk to a lawyer before submitting. Not optional for anything above hobby-project value. This post is not legal advice.

Prepare a repo before it leaves your machine

If you do decide to offer something, don't zip a folder and upload it. Scrub it first. Here's a minimal pre-flight pass you can run locally.

# 1. Scan the full git history for secrets (not just the current tree)
gitleaks detect --source . --log-opts="--all" --report-path leaks.json

# 2. Find hardcoded keys and tokens that scanners miss
grep -rEn "(api[_-]?key|secret|token|passwd|password)\s*[=:]" . \
  --include="*.py" --include="*.js" --include="*.ts" --include="*.yaml" --include="*.env*"

# 3. List every license in your dependency tree
pip-licenses --format=markdown          # Python
npx license-checker --summary           # Node

# 4. Check for client or customer identifiers
grep -rEin "(client_name|acme|customer_id)" . --include="*.py" --include="*.sql"

Then decide what to submit. A cleaned export beats a repo with its history attached:

# Export a clean snapshot with no git history
git archive --format=tar.gz --output=../export-snapshot.tar.gz HEAD

Git history is where leaked credentials hide, so don't submit the .git directory. Rotate any credential that has ever appeared in the history, whether or not you sell anything.

You can also keep a manifest so you know exactly what you offered:

{
  "repo": "internal-invoice-tool",
  "snapshot_date": "2026-10-05",
  "contains_client_data": false,
  "third_party_code_reviewed": true,
  "contractor_assignments_on_file": true,
  "secrets_scan": "clean",
  "approved_by": "owner",
  "license_terms_reviewed_by_counsel": true
}

If you can't truthfully fill in every field, the repo isn't ready.

The bigger leak: what your AI workflows already expose

Here's the part most coverage skips. Selling code is a deliberate, one-time decision with a portal, a price, and a counterparty. The data your daily AI tooling touches is a continuous, mostly invisible flow, and most small teams have never mapped it.

Think about where your code and business data currently go:

  • Coding assistants and IDE plugins that send file contents, diffs, or whole-repo context to a hosted model.
  • Automation tools (n8n, Zapier, Make-style workflows) that pass customer emails, invoices, and CRM records through third-party LLM APIs.
  • Agents with tool access that can read a directory, query a database, or open a ticket, and have no concept of which files are sensitive.
  • Logs and traces. Prompt and response logging often captures the exact data you meant to keep private.
  • Browser extensions and "free" AI utilities with terms nobody read.

Each of these has its own data-handling terms, retention settings, and training defaults. They vary by vendor and plan, and they change. Check each provider's current data-use documentation instead of relying on memory or blog posts, including this one.

A practical way to start is a one-page data-flow inventory:

workflow: lead-followup-agent
data_in:
  - source: CRM export
    contains: [names, emails, deal notes]
    sensitivity: high
model_provider: <provider>
provider_terms_checked: 2026-10-05
training_on_inputs: <verify on current terms page>
retention: <verify>
logged_prompts: true        # <- this is a finding
logs_retention_days: unlimited   # <- so is this
boundaries:
  never_send: [payment data, credentials, contract text]
  redact_before_send: [phone numbers, street addresses]
owner: ops
review_cadence: quarterly

The <verify> fields are deliberate. Don't fill them from assumption. If you can't find the answer on the vendor's current terms, that's your finding.

Set hard boundaries, not good intentions

Policy documents don't stop leaks; code does. Three controls that matter most for small teams:

1. Redact before the call, not after. A thin wrapper that scrubs known-sensitive patterns before any request leaves your network:

import re

PATTERNS = {
    "email": re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+"),
    "api_key": re.compile(r"(?i)(sk|key|token)[-_][A-Za-z0-9]{16,}"),
    "phone": re.compile(r"\+?\d[\d\s().-]{8,}\d"),
}

def scrub(text: str) -> str:
    for label, rx in PATTERNS.items():
        text = rx.sub(f"[REDACTED_{label.upper()}]", text)
    return text

def safe_llm_call(client, prompt: str, **kw):
    return client.generate(scrub(prompt), **kw)

Regex redaction is a floor, not a guarantee. It misses free-text names and context that identifies a client. It beats nothing, and it makes the boundary testable.

2. Allowlist what agents can read. An agent should get access to specific directories and tables, not the whole filesystem. Deny by default, and add paths one at a time with a reason.

3. Separate repos by sensitivity. If you ever want to sell or share a repo, you'll be glad the client work and the secrets weren't mixed into it. Keep customer-bound code in its own repositories with tighter tool access, and keep prototypes and internal utilities somewhere an assistant can freely read.

Log what leaves, too. A request log with destination, size, and redaction count lets you answer "what did we send last month?" in minutes instead of guessing.

A short checklist before you do anything

  1. Inventory: list every place code and customer data flow to a third-party model or service.
  2. Verify terms: read each vendor's current training, retention, and logging terms, and date your notes.
  3. Classify repos: never-sell, caution, candidate.
  4. Confirm ownership: contractor assignments, client contracts, open-source licenses.
  5. Scrub: secrets scan across full history, rotate anything that ever leaked.
  6. Get terms in writing: uses, retention, license scope, deletion, confidentiality, payment.
  7. Get legal review before submitting anything of real value.
  8. Set technical boundaries: redaction, allowlists, logging.

If a buyer can't answer questions 1 and 2 of the terms list in writing, treat that as your answer.

Where this fits in practice

A market for proprietary code makes a question urgent that was already worth asking: what do your AI workflows expose, and who decided that? I'm a senior engineer who builds AI automations that actually ship, so this is the lens I bring to it: before you add another automation or sell another repo, map where the data goes and set the boundaries yourself.

If you're weighing an offer like Google's, or just realized your automations send more than you thought, that boundary-setting is the first job. Reach out through BizFlowAI and we can talk through your current workflows.


Work with BizFlowAI

If you'd rather have this built for you, that's what we do: production AI automation for solo founders and small teams — agents, integrations, and document pipelines that actually ship.

Book a free discovery call — 30 minutes, we map the highest-ROI automation in your workflow. No pitch deck, just engineering.

More guides like this on the BizFlowAI blog.

Frequently asked questions

What is Google's Content Offer Pilot and who can join it?

The Content Offer Pilot is a Google program that pays developers and small businesses for specialized, non-public content, including proprietary code repositories, to improve its products and services. Partners submit selected material through an online portal, choose what to offer, and propose a price. A deal only goes ahead if both sides agree. Google has not published eligibility criteria, though hundreds of partners across roughly 100 countries have reportedly participated.

Does Google use code bought through this program to train AI models like Gemini?

Google has not confirmed this. Its stated purpose is to improve its products and services, and it says it prioritizes science, math and learning contexts. Reports that the code trains Gemini or coding agents are inferences by news outlets, not a Google statement. A June email to Play Store developers said real-world code helps with building coding evals and benchmarks, so check the contract's permitted-uses clause before assuming either way.

What should I check before selling my private code repository to an AI company?

First classify the repo: never sell code with customer data, client-owned IP, secrets or NDA obligations, and be cautious with your core competitive product. Confirm you own every line, including contractor code and open-source licenses. Then get written answers on permitted uses, retention and deletion, license scope (exclusive or not, perpetual or limited), whether outputs can reproduce your code, who can see it, and how payment works. Have a lawyer review the terms before you submit.

How do I scrub a git repo of secrets before sharing it externally?

Run a secrets scanner such as gitleaks across the full git history with the --all log option, since credentials often hide in old commits. Supplement it with grep for hardcoded keys and client identifiers, and audit dependency licenses with pip-licenses or license-checker. Export a clean snapshot with git archive so no .git history leaves your machine. Rotate any credential that has ever appeared in the history, whether or not you sell anything.

If I sell code to Google, do I still own it?

Possibly, but ownership and license rights are different things. A June 2026 email reported by 404 Media described the offer as non-exclusive with developers keeping 100% of their IP. That language is not from a published contract template for the expanded pilot. You can keep ownership while granting broad usage rights, and those rights determine your real exposure, so read the license scope, retention and derivative-use terms carefully.