
When a developer shares an article on Twitter/X, LinkedIn, Discord, Slack, or Bluesky, the link preview card is the single most critical asset for driving click-throughs. If the card shows a generic default banner—or worse, fails to render and displays a raw text link—engagement drops off a cliff.
In modern Next.js applications, the standard way to deliver customized, pixel-perfect social preview cards is through dynamic OpenGraph (OG) image generation powered by @vercel/og (ImageResponse). On paper, it looks trivial: write JSX, pass query parameters, and let the runtime render a crisp 1200x630 PNG.
// A naive dynamic OG image route
import { ImageResponse } from 'next/og';
export async function GET(req: Request) {
const { searchParams } = new URL(req.url);
const title = searchParams.get('title') || 'Default Title';
const authorAvatar = searchParams.get('avatar');
return new ImageResponse(
(
<div style={{ display: 'flex', width: '100%', height: '100%', background: '#000', color: '#fff' }}>
<h1>{title}</h1>
{authorAvatar && <img src={authorAvatar} width={80} height={80} />}
</div>
),
{ width: 1200, height: 630 }
);
}This pattern works deceptively well in local development. But in production on ZyVOP—our developer publishing platform that syndicates technical articles across DEV, Hashnode, Medium, and social channels—this naive pattern collapsed under real-world traffic.
Requests to dynamic OG image endpoints were taking anywhere from 600ms up to 3,800ms. Social crawlers timed out, our server's CPU spiked to 100%, and repeated bot scrapes caused severe request queuing.
Here is the full breakdown of why dynamic OG generation is inherently slow, the anatomy of our latency bottleneck, and how we engineered a two-tier in-memory LRU caching architecture that slashed response times from 3,800ms down to 1.2ms—a 99.8% latency reduction.
1. The Issue: Why Dynamic OG Image Generation Is So Slow
To understand why the route was crawling, we have to look under the hood of what happens when a crawler hits /api/og?title=...&authorImage=....
sequenceDiagram
autonumber
actor Crawler as Social Bot (Twitter/Discord/LinkedIn)
participant Route as Next.js API (/api/og)
participant External as Remote Avatar Host (GitHub/Gravatar/CDN)
participant Satori as Satori (HTML/CSS to SVG Engine)
participant Resvg as Resvg (SVG to PNG Rasterizer)
Crawler->>Route: GET /api/og?title=...&authorImage=https://...
activate Route
Note over Route,External: Latency Trap 1: Network & SSRF Overhead (500ms - 3,000ms)
Route->>External: Outbound HTTP Fetch (DNS + TLS + Download)
External-->>Route: 200 OK (Raw Image Data)
Route->>Route: Magic byte inspection & Base64 encoding
Note over Route,Resvg: Latency Trap 2: CPU-bound Rasterization (200ms - 800ms)
Route->>Satori: Parse JSX + Layout Yoga Engine -> Generate SVG
Satori->>Resvg: Compile SVG -> Rasterize 1200x630 Canvas -> PNG Buffer
Resvg-->>Route: Uint8Array PNG Output (150KB - 250KB)
Route-->>Crawler: 200 OK (image/png) [Total: 1,800ms - 3,800ms]
deactivate RouteThe latency stems from three compounding bottlenecks:
Bottleneck A: The Satori + Resvg CPU Penalty (200ms – 800ms)
ImageResponse relies on Satori (an engine that converts HTML/JSX and CSS flexbox into an SVG) and Resvg (a Rust-based library compiled to WebAssembly that rasterizes the SVG into a binary PNG).
Rasterizing a 1200x630 canvas with multiple typographic layers, gradient borders, badges, SVG icons, and embedded images is pure, intensive CPU work:
Calculating font glyph layout and kerning.
Computing flexbox box models.
Rasterizing vector paths into thousands of pixels.
On a typical cloud VM or container, rendering a single card consumes 200ms to 800ms of dedicated CPU time. If 15 crawlers scrape your links simultaneously, you saturate your CPU cores instantly.
Bottleneck B: The Remote Avatar Network & SSRF Trap (Up to 3,000ms)
A technical article card looks empty without the author's avatar. But author avatars are hosted externally—on GitHub (avatars.githubusercontent.com), Twitter, Gravatar, or user CDNs.
To fetch that image securely, the server must:
Resolve DNS for the external host.
Perform a TCP and TLS handshake.
Download the image bytes over the public internet.
Execute Server-Side Request Forgery (SSRF) defense checks: verifying the host against allowlists, inspecting the initial byte stream for image magic bytes (
0x89 0x50 0x4E 0x47for PNG,0xFF 0xD8 0xFFfor JPEG), validating against chunked compression bombs, and applying anAbortSignal.timeout(3000).
If the author's avatar host has high latency or throttles requests, the OG route stalls for up to 3 full seconds before it can even pass the image to Satori.
Bottleneck C: The Crawler Stampede
When an article is published or shared:
The user posts the link on Twitter/X. Twitterbot hits
/api/og.The link is cross-posted to LinkedIn. LinkedInBot hits
/api/og.Someone pastes the link into a Discord channel. Discordbot hits
/api/og.The link is sent in Slack and iMessage. Slackbot and Applebot hit
/api/og.Search indexers (Googlebot, Bingbot) follow within minutes.
Each bot makes an un-cached, direct GET request. Without application-level memoization, our server was independently executing the exact same remote fetch, the exact same SSRF validation, and the exact same Satori/Resvg rendering 5 to 10 times within seconds for the exact same article.
Bottleneck D: The Self-Hosted Reality
On managed serverless platforms (like Vercel), this inefficiency is masked by spinning up separate Lambda functions for each crawler request—until you look at your invoice and compute-second usage.
When running Next.js self-hosted as a standalone Docker container on a VPS (e.g. 2–4 vCPUs behind Caddy or Nginx), repeated CPU-bound Satori calls block worker threads. The Node.js event loop slows down, and normal users browsing the web app experience degraded response times.
2. Why Common Workarounds Failed
Before designing our solution, we evaluated the standard advice:
"Just put Cloudflare in front of it"
Edge CDN caching is necessary, but it is not sufficient. Cloudflare edge caches are distributed across hundreds of global PoPs (Points of Presence). A request from a Discord bot in Virginia hits an Ashburn PoP; Twitterbot in California hits a San Jose PoP; LinkedIn hits another. Each PoP suffers a cold cache miss and forwards the request to your origin."Pre-generate all OG images at build time"
ZyVOP is an interactive publishing platform where writers publish articles, update titles, edit subtitles, and publish breaking news in real-time. Pre-rendering millions of permutations at static build time is mathematically impossible."Upload generated images to an S3/R2 bucket on save"
While viable, writing to object storage on every draft edit introduces asynchronous orchestration overhead, external database state tracking, and cloud bucket read/write API costs for ephemeral preview cards that may never be shared.
We needed a solution that was:
Zero external dependencies (no new databases, Redis clusters, or bucket queues).
Instantaneous on repeat hits (< 5ms response time).
Safe from out-of-memory (OOM) crashes on containers with limited RAM.
Smart enough to cache both avatar fetches and final PNG renders separately.
3. How We Solved It: Dual-Tier In-Memory LRU Caching
The core architectural realization was that we had two distinct caching problems with two different lifecycles:
The Avatar Lifecycle: Author avatars change very rarely (maybe once every few months), and multiple posts are written by the same author.
The Rendered Image Lifecycle: The final
1200x630PNG is deterministic for a given set of query parameters (title,author,authorImage,date,category,type).
We designed a two-tier in-memory Least Recently Used (LRU) caching architecture using lru-cache.
flowchart TD
Req[Incoming Request: /api/og?title=...&authorImage=...] --> CheckImg{Tier 2: In ogImageCache?}
%% Cache Hit Path
CheckImg -- YES (Cache Hit) --> ServeFast[Return Cached PNG Uint8Array<br/>⚡ 1.2ms | 0% CPU]
%% Cache Miss Path
CheckImg -- NO (Cache Miss) --> CheckAvatar{Tier 1: In avatarCache?}
CheckAvatar -- YES --> UseCachedAv[Use Memory ArrayBuffer<br/>0ms Network]
CheckAvatar -- NO --> FetchAv[fetchSafeOgAvatar: SSRF + HTTP Fetch<br/>~300ms - 2,500ms]
FetchAv --> SaveAv[Store in avatarCache<br/>TTL: 12 Hours]
SaveAv --> UseCachedAv
UseCachedAv --> SatoriRender[Run Satori Flexbox Layout Engine]
SatoriRender --> ResvgRender[Run Resvg Rust/WASM Rasterizer]
ResvgRender --> GenBuf[Generate PNG Buffer]
GenBuf --> SaveImg[Store in ogImageCache<br/>TTL: 24 Hours | Cap: 250MB]
SaveImg --> ServeFresh[Return Fresh PNG Response<br/>⚡ Sets Cache-Control Headers]
style ServeFast fill:#10b981,stroke:#059669,stroke-width:2px,color:#fff
style ServeFresh fill:#3b82f6,stroke:#2563eb,stroke-width:2px,color:#fff
style FetchAv fill:#f59e0b,stroke:#d97706,stroke-width:2px,color:#fff
style ResvgRender fill:#ef4444,stroke:#dc2626,stroke-width:2px,color:#fffTier 1: The Remote Avatar Cache (avatarCache)
Instead of fetching the author's avatar over the network every time a card is generated:
We cache the sanitized, raw
ArrayBufferin memory for 12 hours.Keyed by the
authorImageUrl.If a writer publishes 5 articles in an afternoon, or social bots request 5 different posts from the same author, the network request and SSRF validation run exactly once.
Tier 2: The Rendered OG Image Cache (ogImageCache)
Instead of executing Satori and Resvg on every request:
We cache the final rendered PNG
Bufferin memory for 24 hours.Keyed by the full, normalized query string (
searchParams.toString()).If Twitterbot hits the URL, Satori renders the image. When LinkedInBot, Discordbot, and Slackbot hit the exact same URL seconds later, the server immediately returns the cached binary buffer directly from RAM.
4. The Implementation: Deep Dive into the Code
Let's look at the actual production implementation in our codebase.
Step 1: Building the Shared Cache Module (lib/og-cache.ts)
We created a centralized cache module with strict memory bounds to prevent memory leaks or out-of-memory (OOM) container crashes.
// lib/og-cache.ts
import { LRUCache } from 'lru-cache';
import { fetchSafeOgAvatar } from '@/lib/security/og-avatar';
/**
* In-memory LRU cache for rendered OG image buffers.
* Keyed by the full query string so identical params return the cached PNG
* without re-running Satori + Resvg.
*/
export const ogImageCache = new LRUCache<string, Buffer>({
max: 500,
maxSize: 250_000_000, // Hard memory cap: ~250 MB
sizeCalculation: (value) => value.byteLength,
ttl: 1000 * 60 * 60 * 24, // 24 hours TTL
});
/**
* In-memory LRU cache for remote author avatar image data.
* Eliminates the up-to-3s network round-trip on repeat OG requests
* for the same author.
*
* We use a wrapper object so we can distinguish a "cached failure/empty avatar"
* from a "never fetched" cache miss.
*/
const avatarCache = new LRUCache<string, { data: ArrayBuffer | undefined }>({
max: 200,
ttl: 1000 * 60 * 60 * 12, // 12 hours TTL
});
/**
* Fetch an author avatar with caching. Returns `undefined` when no
* usable image is available (preserving the fetchSafeOgAvatar contract).
*/
export async function getCachedAvatar(
authorImageUrl: string | null | undefined,
): Promise<ArrayBuffer | undefined> {
if (!authorImageUrl) return undefined;
const cached = avatarCache.get(authorImageUrl);
if (cached !== undefined) {
return cached.data;
}
const data = await fetchSafeOgAvatar(authorImageUrl);
avatarCache.set(authorImageUrl, { data });
return data;
}
/**
* Build a high-performance HTTP Response from a cached Buffer.
*/
export function cachedOgResponse(buffer: Buffer): Response {
return new Response(new Uint8Array(buffer), {
headers: {
'Content-Type': 'image/png',
'Cache-Control': 'public, max-age=31536000, immutable',
'X-Cache': 'HIT',
},
});
}Why These Specific Settings Matter:
Memory Budgeting (
maxSize: 250_000_000): A dynamic 1200x630 PNG ranges between 100KB and 250KB. Storing 500 images consumes between 50MB and 125MB of RAM. Even during high-traffic spikes, the cache strictly caps its memory consumption at 250MB, safeguarding our VPS from the Linux OOM killer.The Wrapper Object
{ data: ArrayBuffer | undefined }: Inlru-cachev11+, storingnullorundefineddirectly violates TypeScript's non-nullable object constraint (V extends {}). Wrapping the buffer in an object lets us cache failed avatar fetches (e.g. invalid URL or 404), preventing repeat wasted fetches while satisfying the type system.new Uint8Array(buffer): In modern Node.js and Next.js Route Handlers, passing a raw NodeBufferdirectly tonew Response(buffer)can trigger type mismatches with standardBodyInit. Casting tonew Uint8Array(buffer)guarantees zero-copy compatibility with Web API standards.
Step 2: Integrating the Fast Path into the Route (app/api/og/route.tsx)
In the Next.js API route, we plug in the cache at the very top of the handler. If a match exists, we exit immediately—bypassing GraphQL queries, parameter sanitization, and the entire Satori rendering pipeline.
// app/api/og/route.tsx
import { ImageResponse } from 'next/og';
import { NextRequest } from 'next/server';
import { ogImageCache, getCachedAvatar, cachedOgResponse } from '@/lib/og-cache';
export const runtime = 'nodejs';
export async function GET(req: NextRequest) {
const { searchParams } = new URL(req.url);
// 1. FAST PATH: Return cached PNG buffer immediately on cache hit
const cacheKey = searchParams.toString();
const cached = ogImageCache.get(cacheKey);
if (cached) {
return cachedOgResponse(cached);
}
// Extract parameters
let rawTitle = (searchParams.get('title') || '').trim();
let author = (searchParams.get('author') || '').trim();
let authorImage = searchParams.get('authorImage');
const type = (searchParams.get('type') || '').trim().toLowerCase();
// 2. TIER 1 CACHE: Retrieve author avatar from memory cache
const avatarData = await getCachedAvatar(authorImage);
let avatarSrc: string | undefined;
if (avatarData) {
const u8 = new Uint8Array(avatarData);
const isPng = u8[0] === 0x89 && u8[1] === 0x50 && u8[2] === 0x4e && u8[3] === 0x47;
const mime = isPng ? 'image/png' : 'image/jpeg';
avatarSrc = `data:${mime};base64,${Buffer.from(avatarData).toString('base64')}`;
}
// 3. RENDER WITH SATORI + RESVG
const postResponse = new ImageResponse(
(
<div style={{ /* 1200x630 styles, typography, badges, gradients */ }}>
{avatarSrc && <img src={avatarSrc} width={56} height={56} style={{ borderRadius: '50%' }} />}
<h1>{rawTitle}</h1>
</div>
),
{
width: 1200,
height: 630,
headers: {
'Cache-Control': 'public, max-age=31536000, immutable',
},
},
);
// 4. TIER 2 CACHE: Store rendered PNG buffer in memory for future hits
const postBuffer = Buffer.from(await postResponse.arrayBuffer());
ogImageCache.set(cacheKey, postBuffer);
return cachedOgResponse(postBuffer);
}Step 3: Namespacing Multi-Template Routes (app/api/og/ai/route.tsx)
ZyVOP also features a specialized OG generator for our AI newsletter (/api/og/ai). Because both routes share the same underlying LRU instance, we namespace cache keys to prevent collision:
// app/api/og/ai/route.tsx
export async function GET(req: NextRequest) {
const { searchParams } = new URL(req.url);
// Prefix key with 'ai:' to namespace against the main OG route
const cacheKey = `ai:${searchParams.toString()}`;
const cached = ogImageCache.get(cacheKey);
if (cached) {
return cachedOgResponse(cached);
}
// ... render AI card template ...
const aiBuffer = Buffer.from(await aiResponse.arrayBuffer());
ogImageCache.set(cacheKey, aiBuffer);
return cachedOgResponse(aiBuffer);
}5. The Results: Benchmarking Before and After
We benchmarked the route using an automated suite simulating concurrent crawler traffic with real remote avatars.
Latency Comparison
Scenario | Before (Uncached) | After (Dual-Tier Cache) | Improvement |
|---|---|---|---|
Cold Request (First hit, with remote avatar) | 1,850ms – 3,820ms | 1,850ms (Initial compile & fetch) | Baseline |
Warm Repeat Hit (Twitterbot -> LinkedInBot) | 1,850ms – 3,820ms | 1.2ms | 99.9% faster |
Different post, same author (Avatar hit) | 1,850ms – 3,820ms | 240ms (Zero network fetch) | 87.0% faster |
Brand template (No avatar, repeat hit) | 350ms – 720ms | 0.9ms | 99.7% faster |
Throughput & System Load
Under a benchmark load of 50 concurrent crawler requests:
Before: Server CPU utilization spiked to 100%, average latency rose to 4.2 seconds, and requests began backing up in the queue.
After: Server CPU hovered below 3%, and the route sustained 2,400+ requests per second directly from the Node.js memory buffer.
┌────────────────────────────────────────────────────────┐
│ LATENCY REDUCTION │
├───────────────────────┬────────────────────────────────┤
│ Uncached Post (Worst) │ ████████████████████ 3,800 ms │
│ Uncached Post (Avg) │ ██████████ 1,900 ms │
│ Avatar Hit (No Net) │ █ 240 ms │
│ Warm In-Memory Hit │ ▏ 1.2 ms │
└───────────────────────┴────────────────────────────────┘6. Key Engineering Takeaways
If you are generating dynamic OpenGraph images in Next.js, here are four lessons worth applying:
1. Separate the Avatar Cache from the Image Cache
Never tightly couple remote resource fetching to template rendering. Avatars and post cards have completely different lifecycles. Decoupling them allows an author's avatar to stay warm across their entire catalog of articles.
2. Guard Your Process Memory with maxSize
Never use an unbounded map or simple Set in production. Always use an LRU cache with an explicit maxSize and sizeCalculation: (value) => value.byteLength. 500 images capped at 250MB gives you ample buffer capacity without risking unexpected process termination on a memory-constrained host.
3. Use runtime = 'nodejs' for Heavy In-Memory Caches
While Edge runtimes (V8 isolates) are fast to boot, their memory space is ephemeral and reset frequently between executions. Running your OG route in the nodejs runtime (export const runtime = 'nodejs') allows your in-memory LRU cache to persist across thousands of requests in the long-lived Node.js process.
4. Pair In-Memory Caching with Immutable CDN Headers
The best request is the one that never touches your origin. By attaching Cache-Control: public, max-age=31536000, immutable, downstream CDNs (Cloudflare, Fastly) and social platforms cache the image at the edge for up to a year. The in-memory LRU cache protects your origin against the initial wave of global edge cache misses.
Conclusion
Dynamic social preview cards are essential for modern content syndication, but they shouldn't compromise your application's stability or response times.
By identifying the distinct bottlenecks of Satori CPU rasterization and remote avatar network latency, we transformed a sluggish, 3.8-second endpoint into a sub-2-millisecond microsecond-fast service—using 60 lines of clean TypeScript and zero external infrastructure.
Have questions about dynamic OG generation or developer publishing workflows? Check out our code on GitHub or start publishing with ZyVOP.
Comments (0)
Join the discussion by logging into your account.
No comments yet. Be the first to comment!