P1 incidentINC-2026-02-18Resolvedweb-app

Gemini Image Generation Quota Exhaustion

While reviewing logs on the customer’s infrastructure, ApexData spotted a pattern: the AI image generation feature was failing in daily waves — hundreds of errors clustered in the morning hours, then silence until the next day. The root cause: a third-party API quota was being exhausted within one to two hours of peak traffic. Every request after that point failed silently. Users were seeing generic error messages with no indication that the feature was effectively down for the rest of the day.

Date
2026-02-18
Severity
Medium
Status
Resolved
Affected Service
web-app
Detected By
Audit validation (log and trace analysis)
Client
ACME Corp (EdTech platform)

The Error: "You exceeded your current quota, please check your plan and billing details"

A 429 Too Many Requests with this message from generativelanguage.googleapis.com means a Gemini API quota is exhausted — in this incident the daily cap for one project and model (quotaId GenerateRequestsPerDayPerProjectPerModel), which resets at midnight Pacific time. Until the reset — or a raised plan limit — requests to that model from that project keep failing with the same 429. The fix has two sides: raise the limit or spread the traffic, and stop converting 429 into a generic 500 — classify the upstream status, return a distinct error to the client, and alert on remaining quota before it reaches zero.

Related error: "No image generated"

The same route also failed with Error: No image generated — Gemini accepted the request but returned a response without an image payload. That is not a quota problem: it occurs independently of quota state and needs its own handling — check promptFeedback and finishReason in the response to distinguish blocked generation from a transient empty result, retry only the transient case, and surface a meaningful error instead of a generic 500.

Log Evidence

All 361 error logs originate from the web-app pods. Only 4 ingress-level logs were captured for this endpoint.

429 Quota Exhaustion — Full Error

* Quota exceeded for metric:
    generativelanguage.googleapis.com/generate_requests_per_model_per_day,
    limit: 0
  [{
    "@type": "type.googleapis.com/google.rpc.QuotaFailure",
    "violations": [{
      "quotaMetric": "generativelanguage.googleapis.com/generate_requests_per_model_per_day",
      "quotaId": "GenerateRequestsPerDayPerProjectPerModel"
    }]
  }]

The violation names the daily per-project, per-model quota (GenerateRequestsPerDayPerProjectPerModel) for gemini-2.5-flash-image. Once the cap is hit, every request to the model fails with 429 until the quota resets at midnight Pacific time — 08:00 UTC at this time of year.

429 Quota Exhaustion — Wrapped Error

Error generating image with Gemini: Error: [GoogleGenerativeAI Error]:
Error fetching from https://generativelanguage.googleapis.com/v1beta/
models/gemini-2.5-flash-image:generateContent: [429 Too Many Requests]
You exceeded your current quota, please check your plan and billing details.

Empty Response Error

Error generating image with Gemini: Error: No image generated
    at eo (.next/server/app/ai/generate-image-gemini/route.js:1:21934)

Five occurrences in 25 seconds, suggesting a single user retrying rapidly after each failure.

Traffic Pattern

Hourly Error Breakdown

Hour (UTC)429 Rate LimitEmpty ResponseTotal
Feb 17 00:00055
Feb 17 02:0037037
Feb 17 04:0052052
Feb 17 06:0048048
Feb 17 08:00022
Feb 17 20:00066
Feb 18 06:0025025
Feb 18 07:0066066
Feb 18 12:00088
Feb 18 15:00066

Rate-limit (429) errors cluster between 02:00–08:00 UTC and stop at 08:00 UTC — midnight Pacific, when the daily quota resets. Empty-response errors occur throughout the day independently of quota state. On both days the cap was already exhausted by the early-morning hours, and every attempt kept failing until the reset.

ROOT CAUSE IN MINUTES, NOT DAYS ZERO CODE CHANGES REQUIRED $2.5M SAVED PER 100 DEVS

Root Cause

The Next.js API route calls gemini-2.5-flash-image:generateContent via the @google/generative-ai SDK. Two failure modes:

  • Quota exhaustion (429): The Gemini API key is on a plan with a low request-per-day limit. The SDK throws an error with the 429 status. The route catches this and logs it but returns HTTP 500 to the caller.
  • Empty response: The Gemini API accepts the request but returns no image data. The route throws Error: No image generated and returns HTTP 500.

Code Analysis

// web-app/app/ai/generate-image-gemini/route.ts
} catch (error) {
    console.error("Error generating image with Gemini:", error);
    return NextResponse.json(
        { error: error instanceof Error ? error.message : "Failed to generate image" },
        { status: 500 }  // ← always 500, even for 429
    );
}

The route creates a new GoogleGenerativeAI client on every request (no singleton, no connection pooling). The catch-all error handler returns HTTP 500 for all failures, including 429 rate limits. No retry logic, no rate limiting, no error classification, no circuit breaker.

Impact

  • User-facing: Users attempting to generate AI images get a 500 error. ~690 failed HTTP requests observed in traces over 2 days.
  • No data loss or cascading failures. The feature is self-contained.
  • Revenue impact: If image generation is part of a paid feature flow, users may abandon the flow.

Recommended Remediation

Immediate

  • Return HTTP 429 (not 500) when the Gemini API returns a rate-limit error — parse the error message for "429" and pass retry-after information to the caller
  • Return HTTP 503 with a user-friendly message when the API returns an empty response

Short-term

  • Review the Gemini API plan — upgrade quota or switch to a higher-tier plan if the current limits are insufficient for production traffic
  • Add client-side rate limiting — queue or throttle image generation requests to stay within API quota
  • Add a retry with exponential backoff for transient 429 errors before failing to the user

Medium-term

  • Add a fallback image generation provider (e.g., OpenAI DALL-E, Stability AI) for when Gemini quota is exhausted
  • Add quota monitoring — alert when Gemini API usage reaches 80% of the daily/minute limit so the team can act before users are affected

Want this level of investigation for your infrastructure?

Root cause in minutes, not days.