Architecture Overview
Server Runtime
Rendering Pipeline
Client Navigation
Caching and Export
Development Tools
Build and Configuration
Ecosystem Packages
Testing Infrastructure
How It Works
The following files were used as context for generating this wiki page:
Performance benchmarking provides rigorous automated tooling to measure, track, and compare the developer workflow speed and bundler efficiency of Next.js and Turbopack. It solves the core problem of performance regression by combining repeatable CLI-driven scenario execution with headless browser automation, statistical timing aggregation, and multi-target metric exportation. Key design decisions include robust retry handling for flaky runs, modular interface architecture for reporting results to consoles, JSON outputs, Git snapshots, or Datadog, and integrated tracing plugins for both Webpack and Turbopack. It interacts closely with compiler instrumentation and diagnostic trace servers to provide deep insight into compilation spans, module resolution, and memory usage.
Sources: turbopack/packages/devlow-bench/src/cli.ts:31-57, turbopack/packages/devlow-bench/src/index.ts:47-84, turbopack/packages/devlow-bench/src/runner.ts:29-222, turbopack/packages/devlow-bench/src/browser.ts:21-144, turbopack/packages/devlow-bench/src/interfaces/datadog.ts:31-103, turbopack/packages/webpack-nmt/src/index.ts:32-57, packages/next/src/cli/internal/turbo-trace-server.ts:114-130
The devlow-bench package orchestrates developer workflow benchmarks via a CLI entry point that dispatches execution between run and compare subcommands. It configures scenario variations, validates parameters, and manages the execution runner lifecycle.
The CLI entry point parses process.argv using minimist and detects subcommands from the SUBCOMMANDS set, falling back to legacy run mode for back-compatibility when a bare script path is passed.
process.argv slicing → SUBCOMMANDS check ('run' | 'compare') → runCompareSubcommand / runRunSubcommand → minimist flag parsing → subcommand execution
The runScenarios execution engine processes scenario configurations by expanding Cartesian product permutations of variant properties. It enforces sample counts and warmup runs while handling retries up to MAX_ATTEMPT_MULTIPLIER.
export async function runScenarios(
scenarios: Scenario[],
iface: Interface,
options: { n?: number; warmup?: number } = {}
): Promise<void> {
const n = Math.max(1, Math.floor(options.n ?? 1))
const warmup = Math.max(0, Math.floor(options.warmup ?? 0))
const fullIface = intoFullInterface(iface)
// ... expands variants and iterates through warmup and sample runs ...
}Warning
Do not enable --warmup when measuring cold-start metrics, as discarding the first N runs will invalidate cold-startup timing observations.
The scenario authoring and execution engine manages benchmark definitions, lifecycle event hooks, and statistical timing aggregation. Scenarios are defined via the describe function, which registers configuration matrices and asynchronous runner functions. The execution loop iterates through variant configurations, handles execution attempts with retry boundaries, and computes statistical summaries.
Sources: turbopack/packages/devlow-bench/src/describe.ts:16-63, turbopack/packages/devlow-bench/src/runner.ts:70-201
During a benchmark run, each scenario variant goes through an explicit sequence of interface lifecycle hooks and measurement passes.
runScenarios() → variant iteration → withCurrent() context setup → wrappedIface.start() → 'start' timestamp measurement → variant.scenario.fn() → wrappedIface.end() → summary() statistics aggregation → fullIface.finish()
Sources: turbopack/packages/devlow-bench/src/describe.ts:76-182, turbopack/packages/devlow-bench/src/runner.ts:29-222
Note
During warmup iterations, collecting is set to false, causing all measurements emitted via reportMeasurement or measureTime to be buffered per-run but discarded from final sample aggregation until warmup runs are exhausted.
The FullInterface type defines the complete contract for reporting progress, start and end events, measurements, errors, and variant statistics. The intoFullInterface helper ensures missing optional hooks default to safe asynchronous no-ops.
Warning
Attempt failures inside scenario.fn trigger wrappedIface.error and discard the current attempt's measurements via perRun replacement without interrupting the wider variant sample collection loop, up to MAX_ATTEMPT_MULTIPLIER.
The following example demonstrates how to author a benchmark scenario using describe, capture timing data with measureTime, and report custom metrics via reportMeasurement.
import { describe, measureTime, reportMeasurement, PREVIOUS } from 'devlow-bench'
describe('compiler build pipeline', { mode: ['development', 'production'], incremental: [true, false] }, async (props) => {
// Start an initial timing span
const buildStart = Date.now()
// Simulate compiler compilation task
await Bun.sleep(props.mode === 'development' ? 50 : 200)
// Measure elapsed time relative to the previous timing point or start
await measureTime('build_duration', {
relativeTo: PREVIOUS,
})
// Report a custom numerical memory measurement
await reportMeasurement('peak_memory', 45.2, 'MB', {
props: { incremental: props.incremental },
})
})Important
When relativeTo: PREVIOUS is passed to measureTime or reportMeasurement, the execution engine automatically resolves PREVIOUS to the name of the most recent measurement sharing the same unit string within the current scenario context.
Browser automation and navigation metric capture in devlow-bench is driven by Playwright Chromium sessions managed via the BrowserSession interface and its implementation BrowserSessionImpl. Sessions handle hard navigations, page reloads, and soft interactions like clicks, wrapping executions with automatic resource and performance monitoring.
Sources: turbopack/packages/devlow-bench/src/browser.ts:12-17, turbopack/packages/devlow-bench/src/browser.ts:237-245
The BrowserSessionImpl class coordinates browser instances, contexts, and pages. Hard navigation flows through hardNavigation(metricName, url), which initializes a page if absent, hooks request and console metrics via withRequestMetrics, and records high-resolution timing checkpoints against the /start anchor.
Sources: turbopack/packages/devlow-bench/src/browser.ts:12-17, turbopack/packages/devlow-bench/src/browser.ts:237-285
Warning
Hard navigation throws an error immediately if page.goto produces no response or returns an HTTP response code outside the successful 200..299 range.
The internal networkIdle helper monitors pending network requests by listening to request, requestfailed, and requestfinished events on the Playwright page. It filters out server-sent events (text/event-stream), manages request reference counts, and enforces a grace delay (delayMs = 300) after all requests settle before resolving.
Concurrently, withRequestMetrics intercepts response bodies, categorizes transferred assets by file extension extracted via regular expressions, aggregates response sizes in bytes, tallies request counts, and tracks console message severities (error, warning, log, and uncaught).
const browserSession: BrowserSession = new BrowserSessionImpl(browser, context)
const page = await browserSession.hardNavigation('app_load', 'http://localhost:3000')
await browserSession.softNavigationByClick('nav_click', 'button#load-more')
await browserSession.close()Sources: turbopack/packages/devlow-bench/src/browser.ts:21-144, turbopack/packages/devlow-bench/src/browser.ts:153-235
The devlow-bench reporting subsystem uses modular Interface plugins to export benchmark telemetry across multiple targets. Supported exporter interfaces include console reporting (console), JSON file output (json), Git-backed snapshot tracking (snapshot), historical baseline comparison (compare), and Datadog cloud distribution telemetry (datadog). Each interface implements a subset of lifecycle hooks—such as start, measurement, variantStatistics, error, end, and finish—allowing benchmarks to broadcast metrics to terminals, storage files, version control snapshots, and remote monitoring platforms concurrently.
Sources: turbopack/packages/devlow-bench/src/interfaces/compare.ts:6-32, turbopack/packages/devlow-bench/src/interfaces/datadog.ts:31-103, turbopack/packages/devlow-bench/src/interfaces/console.ts:8-66, turbopack/packages/devlow-bench/src/interfaces/json.ts:18-94, turbopack/packages/devlow-bench/src/interfaces/snapshot.ts:27-68
Each exporter interface is initialized via a factory function that accepts configuration options, environment variables, or file paths.
Sources: turbopack/packages/devlow-bench/src/interfaces/compare.ts:6-32, turbopack/packages/devlow-bench/src/interfaces/datadog.ts:31-103, turbopack/packages/devlow-bench/src/interfaces/console.ts:8-66, turbopack/packages/devlow-bench/src/interfaces/constants.ts:10-24, turbopack/packages/devlow-bench/src/interfaces/json.ts:18-38, turbopack/packages/devlow-bench/src/interfaces/snapshot.ts:10-30
The Datadog exporter constructs a shared set of environment tags during initialization by querying system metadata constants and Git state. The generated tag set includes CI execution status (ci), operating system platform (os) and release (os_release), CPU count (cpus), CPU model (cpu_model), current username (user), CPU architecture (arch), total system memory rounded to gigabytes (total_memory), Node.js version (node_version), and Git commit SHA (git_sha) and branch (git_branch).
Incoming metric names are normalized via toIdentifier, replacing slashes with dots and spaces with underscores. Unit strings such as ms, requests, and bytes are mapped to Datadog metadata unit types (millisecond, request, and byte). During measurement reporting, distribution points are accumulated in memory and submitted in bulk when end is invoked.
const datadogIface = datadogInterface({
apiKey: process.env.DATADOG_API_KEY,
appKey: process.env.DATADOG_APP_KEY,
host: 'ci-runner-01',
})Sources: turbopack/packages/devlow-bench/src/interfaces/datadog.ts:21-103, turbopack/packages/devlow-bench/src/interfaces/constants.ts:1-34
Warning
The Datadog interface throws an immediate runtime error during initialization if DATADOG_API_KEY is missing from the environment or options.
The snapshot interface collects sample rows during variantStatistics execution, tagging each individual sample with its index, value, unit, and relative baseline reference. When finish runs, it calls readGitInfo to resolve the current Git SHA and branch (falling back to GITHUB_SHA and GITHUB_REF_NAME environment variables or executing git rev-parse via child processes), updates all accumulated rows with these Git identifiers, and writes them out using writeSnapshot.
Conversely, the compare interface reads an existing baseline snapshot path via readSnapshot, groups rows using groupRows, and registers variant statistics into a current map. When the benchmark finishes, printComparison evaluates differences between the baseline and current metrics.
Sources: turbopack/packages/devlow-bench/src/interfaces/compare.ts:6-32, turbopack/packages/devlow-bench/src/interfaces/snapshot.ts:10-68
Note
When n === 1, the JSON exporter outputs individual measurements with a single-run schema containing key, value, unit, text, datapoints: 1, and relativeTo. When n > 1, it wraps an array of computed statistical aggregates including mean, p50, p90, and the raw samples array inside a { results } payload object.
Compiler instrumentation plugins track node module resolution and execution performance during builds. The NodeModuleTracePlugin and createNodeFileTrace integrate with Webpack and Next.js compilation pipelines to execute @vercel/experimental-nft file tracing on output chunks. Additionally, trace span utilities provide granular timing measurement and hierarchical telemetry for asynchronous operations.
Sources: turbopack/packages/webpack-nmt/src/index.ts:6-140, turbopack/packages/turbo-tracing-next-plugin/src/index.ts:1-27, packages/next/src/trace/trace.ts:30-154
The NodeModuleTracePlugin accepts configuration options through NodeModuleTracePluginOptions. During the compilation lifecycle, apply() hooks into compiler.hooks.compilation to tap into compilation.hooks.processAssets at stage Compilation.PROCESS_ASSETS_STAGE_SUMMARIZE, triggering createTraceAssets() which inspects entrypoints and populates chunksToTrace.
Warning
Files ending in .wasm or .map are explicitly filtered out by isTraceable and will never be added to chunksToTrace.
Once compilation emits files, compiler.hooks.afterEmit invokes runTrace(), which executes the node-file-trace binary using chunk batching.
The tracing invocation lifecycle proceeds as follows: runTrace() resolves binary paths via require.resolve() for @vercel/experimental-nft/package.json and its platform-specific binary package @vercel/experimental-nft-${process.platform}-${process.arch}/package.json → traceChunks() spawns the node-file-trace child process with spawn('node-file-trace', [...args, ...chunks]) → event listeners pipe stdout and stderr to process streams → exit event resolves or rejects the execution promise.
import { createNodeFileTrace } from '@vercel/turbo-tracing-next-plugin'
const withTrace = createNodeFileTrace({
maxFiles: 64,
log: { level: 'info', detail: true }
})Internal diagnostic management provides HTTP trace server visualization, memory profiling summaries, and blob storage upload tooling to diagnose performance bottlenecks and runtime memory usage. The trace server tooling interacts with native SWC bindings and telemetry spans to inspect trace files, format execution durations, summarize TurboMalloc live bytes, and upload performance profiles securely.
Sources: packages/next/src/cli/internal/turbo-trace-server.ts:1-122, packages/next/src/cli/internal/upload-trace.ts:1-116
The trace server maps internal ticks to human-readable units where 100 internal ticks equal 1 microsecond. Durations are formatted via formatDuration() into microseconds (µs), milliseconds (ms), or seconds (s), while relative timings handle negative offsets when child spans precede their parent reference point. Memory consumption is tracked via summarizeMemorySamples() using TurboMalloc live sample arrays.
Note
The trace server requires native non-WASM SWC bindings; attempting to start the server when native bindings fail to load writes an error message to console.error and terminates the process with exit code 1.
Diagnostic profiles and trace files generated during builds can be uploaded to remote storage via uploadTraceToBlob(). The tool scans the .next-profiles directory for .cpuprofile and trace-turbopack.bin files, validates file headers, requests authorization tokens, and streams contents with progress feedback.
The trace upload execution sequence proceeds as follows: uploadTraceToBlob() reads the .next-profiles directory via fs.readdir() → filters entries matching .cpuprofile or trace-turbopack.bin → reads file headers into a 16-byte buffer via fs.open() and fd.read() → validates V8 CPU profile headers ({"nodes":) or Turbopack trace headers (TRACEv0) via validateCpuProfile() or validateTurbopackTrace() → sends a POST request to getUploadUrl() to acquire an upload token → executes private bucket upload via put() with chunked progress streams or full buffers.
import { uploadTraceToBlob } from 'next/dist/cli/internal/upload-trace'
await uploadTraceToBlob({
directory: process.cwd(),
})