Architecture

From 3,800ms to 1.2ms: Optimizing Dynamic OpenGraph Image Generation in Next.js

How we tackled CPU-heavy Satori rasterization, SSRF-protected avatar latency, and memory constraints using dual-tier LRU caching.

Sanju Singh
•
11 min read
From 3,800ms to 1.2ms: Optimizing Dynamic OpenGraph Image Generation in Next.js

When a developer shares an article on Twitter/X, LinkedIn, Discord, Slack, or Bluesky, the link preview card is the single most critical asset for driving click-throughs. If the card shows a generic default banner—or worse, fails to render and displays a raw text link—engagement drops off a cliff.

In modern Next.js applications, the standard way to deliver customized, pixel-perfect social preview cards is through dynamic OpenGraph (OG) image generation powered by @vercel/og (ImageResponse). On paper, it looks trivial: write JSX, pass query parameters, and let the runtime render a crisp 1200x630 PNG.

// A naive dynamic OG image route
import { ImageResponse } from 'next/og';

export async function GET(req: Request) {
  const { searchParams } = new URL(req.url);
  const title = searchParams.get('title') || 'Default Title';
  const authorAvatar = searchParams.get('avatar');

  return new ImageResponse(
    (
      <div style={{ display: 'flex', width: '100%', height: '100%', background: '#000', color: '#fff' }}>
        <h1>{title}</h1>
        {authorAvatar && <img src={authorAvatar} width={80} height={80} />}
      </div>
    ),
    { width: 1200, height: 630 }
  );
}

This pattern works deceptively well in local development. But in production on ZyVOP—our developer publishing platform that syndicates technical articles across DEV, Hashnode, Medium, and social channels—this naive pattern collapsed under real-world traffic.

Requests to dynamic OG image endpoints were taking anywhere from 600ms up to 3,800ms. Social crawlers timed out, our server's CPU spiked to 100%, and repeated bot scrapes caused severe request queuing.

Here is the full breakdown of why dynamic OG generation is inherently slow, the anatomy of our latency bottleneck, and how we engineered a two-tier in-memory LRU caching architecture that slashed response times from 3,800ms down to 1.2ms—a 99.8% latency reduction.


1. The Issue: Why Dynamic OG Image Generation Is So Slow

To understand why the route was crawling, we have to look under the hood of what happens when a crawler hits /api/og?title=...&authorImage=....

sequenceDiagram
    autonumber
    actor Crawler as Social Bot (Twitter/Discord/LinkedIn)
    participant Route as Next.js API (/api/og)
    participant External as Remote Avatar Host (GitHub/Gravatar/CDN)
    participant Satori as Satori (HTML/CSS to SVG Engine)
    participant Resvg as Resvg (SVG to PNG Rasterizer)

    Crawler->>Route: GET /api/og?title=...&authorImage=https://...
    activate Route
    Note over Route,External: Latency Trap 1: Network & SSRF Overhead (500ms - 3,000ms)
    Route->>External: Outbound HTTP Fetch (DNS + TLS + Download)
    External-->>Route: 200 OK (Raw Image Data)
    Route->>Route: Magic byte inspection & Base64 encoding
    
    Note over Route,Resvg: Latency Trap 2: CPU-bound Rasterization (200ms - 800ms)
    Route->>Satori: Parse JSX + Layout Yoga Engine -> Generate SVG
    Satori->>Resvg: Compile SVG -> Rasterize 1200x630 Canvas -> PNG Buffer
    Resvg-->>Route: Uint8Array PNG Output (150KB - 250KB)
    
    Route-->>Crawler: 200 OK (image/png) [Total: 1,800ms - 3,800ms]
    deactivate Route

The latency stems from three compounding bottlenecks:

Bottleneck A: The Satori + Resvg CPU Penalty (200ms – 800ms)

ImageResponse relies on Satori (an engine that converts HTML/JSX and CSS flexbox into an SVG) and Resvg (a Rust-based library compiled to WebAssembly that rasterizes the SVG into a binary PNG).

Rasterizing a 1200x630 canvas with multiple typographic layers, gradient borders, badges, SVG icons, and embedded images is pure, intensive CPU work:

  • Calculating font glyph layout and kerning.

  • Computing flexbox box models.

  • Rasterizing vector paths into thousands of pixels.

On a typical cloud VM or container, rendering a single card consumes 200ms to 800ms of dedicated CPU time. If 15 crawlers scrape your links simultaneously, you saturate your CPU cores instantly.

Bottleneck B: The Remote Avatar Network & SSRF Trap (Up to 3,000ms)

A technical article card looks empty without the author's avatar. But author avatars are hosted externally—on GitHub (avatars.githubusercontent.com), Twitter, Gravatar, or user CDNs.

To fetch that image securely, the server must:

  1. Resolve DNS for the external host.

  2. Perform a TCP and TLS handshake.

  3. Download the image bytes over the public internet.

  4. Execute Server-Side Request Forgery (SSRF) defense checks: verifying the host against allowlists, inspecting the initial byte stream for image magic bytes (0x89 0x50 0x4E 0x47 for PNG, 0xFF 0xD8 0xFF for JPEG), validating against chunked compression bombs, and applying an AbortSignal.timeout(3000).

If the author's avatar host has high latency or throttles requests, the OG route stalls for up to 3 full seconds before it can even pass the image to Satori.

Bottleneck C: The Crawler Stampede

When an article is published or shared:

  • The user posts the link on Twitter/X. Twitterbot hits /api/og.

  • The link is cross-posted to LinkedIn. LinkedInBot hits /api/og.

  • Someone pastes the link into a Discord channel. Discordbot hits /api/og.

  • The link is sent in Slack and iMessage. Slackbot and Applebot hit /api/og.

  • Search indexers (Googlebot, Bingbot) follow within minutes.

Each bot makes an un-cached, direct GET request. Without application-level memoization, our server was independently executing the exact same remote fetch, the exact same SSRF validation, and the exact same Satori/Resvg rendering 5 to 10 times within seconds for the exact same article.

Bottleneck D: The Self-Hosted Reality

On managed serverless platforms (like Vercel), this inefficiency is masked by spinning up separate Lambda functions for each crawler request—until you look at your invoice and compute-second usage.

When running Next.js self-hosted as a standalone Docker container on a VPS (e.g. 2–4 vCPUs behind Caddy or Nginx), repeated CPU-bound Satori calls block worker threads. The Node.js event loop slows down, and normal users browsing the web app experience degraded response times.


2. Why Common Workarounds Failed

Before designing our solution, we evaluated the standard advice:

  1. "Just put Cloudflare in front of it"
    Edge CDN caching is necessary, but it is not sufficient. Cloudflare edge caches are distributed across hundreds of global PoPs (Points of Presence). A request from a Discord bot in Virginia hits an Ashburn PoP; Twitterbot in California hits a San Jose PoP; LinkedIn hits another. Each PoP suffers a cold cache miss and forwards the request to your origin.

  2. "Pre-generate all OG images at build time"
    ZyVOP is an interactive publishing platform where writers publish articles, update titles, edit subtitles, and publish breaking news in real-time. Pre-rendering millions of permutations at static build time is mathematically impossible.

  3. "Upload generated images to an S3/R2 bucket on save"
    While viable, writing to object storage on every draft edit introduces asynchronous orchestration overhead, external database state tracking, and cloud bucket read/write API costs for ephemeral preview cards that may never be shared.

We needed a solution that was:

  • Zero external dependencies (no new databases, Redis clusters, or bucket queues).

  • Instantaneous on repeat hits (< 5ms response time).

  • Safe from out-of-memory (OOM) crashes on containers with limited RAM.

  • Smart enough to cache both avatar fetches and final PNG renders separately.


3. How We Solved It: Dual-Tier In-Memory LRU Caching

The core architectural realization was that we had two distinct caching problems with two different lifecycles:

  1. The Avatar Lifecycle: Author avatars change very rarely (maybe once every few months), and multiple posts are written by the same author.

  2. The Rendered Image Lifecycle: The final 1200x630 PNG is deterministic for a given set of query parameters (title, author, authorImage, date, category, type).

We designed a two-tier in-memory Least Recently Used (LRU) caching architecture using lru-cache.

flowchart TD
    Req[Incoming Request: /api/og?title=...&authorImage=...] --> CheckImg{Tier 2: In ogImageCache?}
    
    %% Cache Hit Path
    CheckImg -- YES (Cache Hit) --> ServeFast[Return Cached PNG Uint8Array<br/>⚡ 1.2ms | 0% CPU]
    
    %% Cache Miss Path
    CheckImg -- NO (Cache Miss) --> CheckAvatar{Tier 1: In avatarCache?}
    
    CheckAvatar -- YES --> UseCachedAv[Use Memory ArrayBuffer<br/>0ms Network]
    CheckAvatar -- NO --> FetchAv[fetchSafeOgAvatar: SSRF + HTTP Fetch<br/>~300ms - 2,500ms]
    FetchAv --> SaveAv[Store in avatarCache<br/>TTL: 12 Hours]
    SaveAv --> UseCachedAv
    
    UseCachedAv --> SatoriRender[Run Satori Flexbox Layout Engine]
    SatoriRender --> ResvgRender[Run Resvg Rust/WASM Rasterizer]
    ResvgRender --> GenBuf[Generate PNG Buffer]
    
    GenBuf --> SaveImg[Store in ogImageCache<br/>TTL: 24 Hours | Cap: 250MB]
    SaveImg --> ServeFresh[Return Fresh PNG Response<br/>⚡ Sets Cache-Control Headers]

    style ServeFast fill:#10b981,stroke:#059669,stroke-width:2px,color:#fff
    style ServeFresh fill:#3b82f6,stroke:#2563eb,stroke-width:2px,color:#fff
    style FetchAv fill:#f59e0b,stroke:#d97706,stroke-width:2px,color:#fff
    style ResvgRender fill:#ef4444,stroke:#dc2626,stroke-width:2px,color:#fff

Tier 1: The Remote Avatar Cache (avatarCache)

Instead of fetching the author's avatar over the network every time a card is generated:

  • We cache the sanitized, raw ArrayBuffer in memory for 12 hours.

  • Keyed by the authorImageUrl.

  • If a writer publishes 5 articles in an afternoon, or social bots request 5 different posts from the same author, the network request and SSRF validation run exactly once.

Tier 2: The Rendered OG Image Cache (ogImageCache)

Instead of executing Satori and Resvg on every request:

  • We cache the final rendered PNG Buffer in memory for 24 hours.

  • Keyed by the full, normalized query string (searchParams.toString()).

  • If Twitterbot hits the URL, Satori renders the image. When LinkedInBot, Discordbot, and Slackbot hit the exact same URL seconds later, the server immediately returns the cached binary buffer directly from RAM.


4. The Implementation: Deep Dive into the Code

Let's look at the actual production implementation in our codebase.

Step 1: Building the Shared Cache Module (lib/og-cache.ts)

We created a centralized cache module with strict memory bounds to prevent memory leaks or out-of-memory (OOM) container crashes.

// lib/og-cache.ts
import { LRUCache } from 'lru-cache';
import { fetchSafeOgAvatar } from '@/lib/security/og-avatar';

/**
 * In-memory LRU cache for rendered OG image buffers.
 * Keyed by the full query string so identical params return the cached PNG
 * without re-running Satori + Resvg.
 */
export const ogImageCache = new LRUCache<string, Buffer>({
  max: 500,
  maxSize: 250_000_000, // Hard memory cap: ~250 MB
  sizeCalculation: (value) => value.byteLength,
  ttl: 1000 * 60 * 60 * 24, // 24 hours TTL
});

/**
 * In-memory LRU cache for remote author avatar image data.
 * Eliminates the up-to-3s network round-trip on repeat OG requests
 * for the same author.
 *
 * We use a wrapper object so we can distinguish a "cached failure/empty avatar"
 * from a "never fetched" cache miss.
 */
const avatarCache = new LRUCache<string, { data: ArrayBuffer | undefined }>({
  max: 200,
  ttl: 1000 * 60 * 60 * 12, // 12 hours TTL
});

/**
 * Fetch an author avatar with caching. Returns `undefined` when no
 * usable image is available (preserving the fetchSafeOgAvatar contract).
 */
export async function getCachedAvatar(
  authorImageUrl: string | null | undefined,
): Promise<ArrayBuffer | undefined> {
  if (!authorImageUrl) return undefined;

  const cached = avatarCache.get(authorImageUrl);
  if (cached !== undefined) {
    return cached.data;
  }

  const data = await fetchSafeOgAvatar(authorImageUrl);
  avatarCache.set(authorImageUrl, { data });
  return data;
}

/**
 * Build a high-performance HTTP Response from a cached Buffer.
 */
export function cachedOgResponse(buffer: Buffer): Response {
  return new Response(new Uint8Array(buffer), {
    headers: {
      'Content-Type': 'image/png',
      'Cache-Control': 'public, max-age=31536000, immutable',
      'X-Cache': 'HIT',
    },
  });
}

Why These Specific Settings Matter:

  1. Memory Budgeting (maxSize: 250_000_000): A dynamic 1200x630 PNG ranges between 100KB and 250KB. Storing 500 images consumes between 50MB and 125MB of RAM. Even during high-traffic spikes, the cache strictly caps its memory consumption at 250MB, safeguarding our VPS from the Linux OOM killer.

  2. The Wrapper Object { data: ArrayBuffer | undefined }: In lru-cache v11+, storing null or undefined directly violates TypeScript's non-nullable object constraint (V extends {}). Wrapping the buffer in an object lets us cache failed avatar fetches (e.g. invalid URL or 404), preventing repeat wasted fetches while satisfying the type system.

  3. new Uint8Array(buffer): In modern Node.js and Next.js Route Handlers, passing a raw Node Buffer directly to new Response(buffer) can trigger type mismatches with standard BodyInit. Casting to new Uint8Array(buffer) guarantees zero-copy compatibility with Web API standards.


Step 2: Integrating the Fast Path into the Route (app/api/og/route.tsx)

In the Next.js API route, we plug in the cache at the very top of the handler. If a match exists, we exit immediately—bypassing GraphQL queries, parameter sanitization, and the entire Satori rendering pipeline.

// app/api/og/route.tsx
import { ImageResponse } from 'next/og';
import { NextRequest } from 'next/server';
import { ogImageCache, getCachedAvatar, cachedOgResponse } from '@/lib/og-cache';

export const runtime = 'nodejs';

export async function GET(req: NextRequest) {
  const { searchParams } = new URL(req.url);

  // 1. FAST PATH: Return cached PNG buffer immediately on cache hit
  const cacheKey = searchParams.toString();
  const cached = ogImageCache.get(cacheKey);
  if (cached) {
    return cachedOgResponse(cached);
  }

  // Extract parameters
  let rawTitle = (searchParams.get('title') || '').trim();
  let author = (searchParams.get('author') || '').trim();
  let authorImage = searchParams.get('authorImage');
  const type = (searchParams.get('type') || '').trim().toLowerCase();

  // 2. TIER 1 CACHE: Retrieve author avatar from memory cache
  const avatarData = await getCachedAvatar(authorImage);
  let avatarSrc: string | undefined;
  if (avatarData) {
    const u8 = new Uint8Array(avatarData);
    const isPng = u8[0] === 0x89 && u8[1] === 0x50 && u8[2] === 0x4e && u8[3] === 0x47;
    const mime = isPng ? 'image/png' : 'image/jpeg';
    avatarSrc = `data:${mime};base64,${Buffer.from(avatarData).toString('base64')}`;
  }

  // 3. RENDER WITH SATORI + RESVG
  const postResponse = new ImageResponse(
    (
      <div style={{ /* 1200x630 styles, typography, badges, gradients */ }}>
        {avatarSrc && <img src={avatarSrc} width={56} height={56} style={{ borderRadius: '50%' }} />}
        <h1>{rawTitle}</h1>
      </div>
    ),
    {
      width: 1200,
      height: 630,
      headers: {
        'Cache-Control': 'public, max-age=31536000, immutable',
      },
    },
  );

  // 4. TIER 2 CACHE: Store rendered PNG buffer in memory for future hits
  const postBuffer = Buffer.from(await postResponse.arrayBuffer());
  ogImageCache.set(cacheKey, postBuffer);

  return cachedOgResponse(postBuffer);
}

Step 3: Namespacing Multi-Template Routes (app/api/og/ai/route.tsx)

ZyVOP also features a specialized OG generator for our AI newsletter (/api/og/ai). Because both routes share the same underlying LRU instance, we namespace cache keys to prevent collision:

// app/api/og/ai/route.tsx
export async function GET(req: NextRequest) {
  const { searchParams } = new URL(req.url);

  // Prefix key with 'ai:' to namespace against the main OG route
  const cacheKey = `ai:${searchParams.toString()}`;
  const cached = ogImageCache.get(cacheKey);
  if (cached) {
    return cachedOgResponse(cached);
  }

  // ... render AI card template ...
  const aiBuffer = Buffer.from(await aiResponse.arrayBuffer());
  ogImageCache.set(cacheKey, aiBuffer);
  return cachedOgResponse(aiBuffer);
}

5. The Results: Benchmarking Before and After

We benchmarked the route using an automated suite simulating concurrent crawler traffic with real remote avatars.

Latency Comparison

Scenario

Before (Uncached)

After (Dual-Tier Cache)

Improvement

Cold Request (First hit, with remote avatar)

1,850ms – 3,820ms

1,850ms (Initial compile & fetch)

Baseline

Warm Repeat Hit (Twitterbot -> LinkedInBot)

1,850ms – 3,820ms

1.2ms

99.9% faster

Different post, same author (Avatar hit)

1,850ms – 3,820ms

240ms (Zero network fetch)

87.0% faster

Brand template (No avatar, repeat hit)

350ms – 720ms

0.9ms

99.7% faster

Throughput & System Load

Under a benchmark load of 50 concurrent crawler requests:

  • Before: Server CPU utilization spiked to 100%, average latency rose to 4.2 seconds, and requests began backing up in the queue.

  • After: Server CPU hovered below 3%, and the route sustained 2,400+ requests per second directly from the Node.js memory buffer.

┌────────────────────────────────────────────────────────┐
│                   LATENCY REDUCTION                    │
├───────────────────────┬────────────────────────────────┤
│ Uncached Post (Worst) │ ████████████████████ 3,800 ms  │
│ Uncached Post (Avg)   │ ██████████ 1,900 ms            │
│ Avatar Hit (No Net)   │ █ 240 ms                       │
│ Warm In-Memory Hit    │ ▏ 1.2 ms                       │
└───────────────────────┴────────────────────────────────┘

6. Key Engineering Takeaways

If you are generating dynamic OpenGraph images in Next.js, here are four lessons worth applying:

1. Separate the Avatar Cache from the Image Cache

Never tightly couple remote resource fetching to template rendering. Avatars and post cards have completely different lifecycles. Decoupling them allows an author's avatar to stay warm across their entire catalog of articles.

2. Guard Your Process Memory with maxSize

Never use an unbounded map or simple Set in production. Always use an LRU cache with an explicit maxSize and sizeCalculation: (value) => value.byteLength. 500 images capped at 250MB gives you ample buffer capacity without risking unexpected process termination on a memory-constrained host.

3. Use runtime = 'nodejs' for Heavy In-Memory Caches

While Edge runtimes (V8 isolates) are fast to boot, their memory space is ephemeral and reset frequently between executions. Running your OG route in the nodejs runtime (export const runtime = 'nodejs') allows your in-memory LRU cache to persist across thousands of requests in the long-lived Node.js process.

4. Pair In-Memory Caching with Immutable CDN Headers

The best request is the one that never touches your origin. By attaching Cache-Control: public, max-age=31536000, immutable, downstream CDNs (Cloudflare, Fastly) and social platforms cache the image at the edge for up to a year. The in-memory LRU cache protects your origin against the initial wave of global edge cache misses.


Conclusion

Dynamic social preview cards are essential for modern content syndication, but they shouldn't compromise your application's stability or response times.

By identifying the distinct bottlenecks of Satori CPU rasterization and remote avatar network latency, we transformed a sluggish, 3.8-second endpoint into a sub-2-millisecond microsecond-fast service—using 60 lines of clean TypeScript and zero external infrastructure.

Have questions about dynamic OG generation or developer publishing workflows? Check out our code on GitHub or start publishing with ZyVOP.

Comments (0)

Join the discussion by logging into your account.

No comments yet. Be the first to comment!

Sanju Singh

Passionate developer sharing knowledge about modern web technologies and best practices.

Subscribe to Sanju Singh's Newsletter

Direct email dispatches when new stories are published. Zero algorithms.