Find and fix game performance problems methodically — measure with the engine profiler first, reason about the frame-time budget, locate the CPU-vs-GPU…
Performance optimization
Performance work is a measurement discipline, not a bag of tricks. The method is always the
same: profile → find the one bottleneck → fix that → measure again. This skill teaches that
loop and the highest-leverage fixes (pooling, batching, allocation control, asset budgets), and
points you at each engine's profiler. It pairs with physics-tuning for simulation cost.
When to use
Use when the frame rate is low or uneven, the game stutters/hitches, or it must hit a target
(60 FPS desktop, 30/60 mobile) and currently doesn't.
Use to decide what to optimize: profile, read the frame budget, and identify whether the CPU
or GPU is the bottleneck before changing any code.
Use to apply specific fixes: object pooling, draw-call/batch reduction, removing per-frame
allocations and GC spikes, and setting asset budgets.
When not to use: for physics jitter/tunneling/timestep specifically, use physics-tuning.
For the engine's concrete profiler UI and rendering settings, use that
engine skill (godot-export covers some build settings; engine cores cover the rest). This skill
is the cross-engine method and the shared fixes.
The golden rule: measure first, never guess
Most performance "fixes" applied without profiling target the wrong thing and add complexity for
no gain. Do not optimize code you have not measured. Open the profiler, find the single
biggest cost in a representative scene on representative hardware, and fix that. Re-measure to
confirm the fix helped before moving on. Profile a release/optimized build where it matters —
editor and debug builds lie (editor overhead, no compiler optimization).
Core workflow
Define the target and reproduce. State the goal (e.g. 60 FPS = 16.67 ms/frame) and find a
repeatable worst-case scene. "Sometimes slow" is unfixable; a reproducible spike is fixable.
Profile before touching code. Run the engine profiler and read the frame: total frame
time, and the split between CPU (game logic, physics, scripts) and GPU (rendering).
Find the bottleneck — CPU or GPU. If GPU time ≫ CPU, attack draw calls/overdraw/shaders/
resolution. If CPU time dominates, attack scripts/physics/allocations. Fixing the wrong side
does nothing.
Fix the single biggest cost. Prefer an algorithmic win (do less work, cache, spatial
partition, run less often) over micro-optimizing a hot line. Apply the matching shared fix
(pooling, batching, allocation removal).
Re-measure on the same scene/hardware. Confirm the number moved. Keep or revert based on
data, not intuition.
Set budgets so it stays fixed. Per-frame ms budgets per subsystem, plus asset budgets
(texture sizes, triangle counts, draw-call ceilings); add a perf check to verification.
Report measured numbers. State before/after frame time, the bottleneck found, and the fix
— never "should be faster". If you could only measure in-editor, say so.
Patterns
1. Frame budget math (turn "feels slow" into a number)
target FPS → frame budget: 60 FPS = 16.67 ms | 30 FPS = 33.3 ms | 120 FPS = 8.33 ms
The WHOLE frame (CPU sim + render submit + GPU) must fit the budget; the GPU runs in parallel,
so the slower of CPU-frame and GPU-frame sets your FPS. Allocate sub-budgets, e.g. @60 FPS:
gameplay/scripts ~5 ms · physics ~3 ms · rendering(CPU submit) ~4 ms · UI/other ~2 ms · slack.
If one subsystem blows its slice, that's your target — not whatever you assumed.
2. Measure with the engine profiler (do this before any fix)
Godot 4.7 : Debugger ▸ Profiler (script/physics time) and Monitors tab (FPS, draw calls, memory).
In code: Performance.get_monitor(Performance.TIME_PROCESS) and
Performance.get_monitor(Performance.RENDER_TOTAL_DRAW_CALLS_IN_FRAME).
Unity 6.3 LTS : Profiler window (CPU/GPU/Memory/Rendering modules) + Frame Debugger for draw calls.
In code: a ProfilerRecorder tracking "CPU Main Thread Frame Time" for a HUD/log.
Unreal 5 : `stat unit` (Frame/Game/Draw/GPU ms), `stat fps`, `stat scenerendering` (draw calls);
Unreal Insights for deep traces.
# Read the split: is the Draw/GPU line the biggest, or the Game/CPU line? That decides the fix.
3. Object pooling (stop allocating/freeing in hot loops)
# Bullets, particles, enemies, damage numbers: reuse a fixed set instead of instantiate()/free()
# every frame — that thrashes memory and (in C#) feeds the GC.
var _pool: Array[Node] = []
func acquire() -> Node:
var n: Node = _pool.pop_back() if not _pool.is_empty() else bullet_scene.instantiate()
n.set_process(true); n.visible = true
return n
func release(n: Node) -> void:
n.set_process(false); n.visible = false # disable + hide; DON'T free
_pool.append(n) # back to the pool for reuse
# RIGHT: pre-warm the pool at load; reuse. WRONG: instantiate()/queue_free() per shot.
4. Cut draw calls (the most common GPU-side win)
Each unique material/texture/state change is roughly a draw call; thousands of them stall the GPU.
- Atlas textures and share materials so sprites/meshes batch into one call.
- Identical meshes → GPU instancing (Unity), MultiMesh / MultiMeshInstance (Godot), Instanced
Static Mesh (Unreal).
- Static geometry → static batching / baking; mark non-moving objects static.
- Reduce overdraw: limit large overlapping transparent/particle layers (they re-shade pixels).
- Fewer real-time lights/shadows; bake lighting where it doesn't move.
Measure draw calls before and after — the count should drop, and so should GPU frame time.
5. Kill per-frame allocations (GC spikes = stutter)
// Unity 6.3 LTS (C#). Allocating every frame fills the managed heap; the GC then stalls a frame.
// WRONG (allocates each call): foreach (var e in FindObjectsOfType<Enemy>()) ... // + LINQ, new[]
// RIGHT: cache references once, reuse buffers, avoid LINQ/boxing in Update.
void Update() {
_hits = Physics.RaycastNonAlloc(ray, _hitBuffer); // reuse a preallocated array
for (int i = 0; i < _hits; i++) { /* ... */ } // no per-frame allocation
}
// Godot/GDScript: avoid building new arrays/dictionaries every frame in _process; reuse them.
Pitfalls
Optimizing without profiling. The intuitive culprit is usually wrong. Measure first, every
time.
Profiling the editor / a debug build. Editor overhead and unoptimized code mislead. Profile
a release build on target hardware for real numbers.
Fixing the wrong side. Micro-optimizing CPU code when the GPU is the bottleneck (or vice
versa) changes nothing. Check the CPU-vs-GPU split first.
Micro-optimizing over algorithm. Shaving a function when an O(n²) loop or a per-frame
full-scene query is the real cost. Reduce the work, don't polish it.
Instantiate/free in hot loops. Spawning and destroying bullets/particles every frame causes
fragmentation and GC spikes. Pool them.
Per-frame allocations / LINQ / boxing in Update (C#) feed the GC → periodic hitches.
Cache and reuse.
Draw-call explosion from unique materials and unbatched sprites/meshes. Atlas, share
materials, instance, batch.
Overdraw from stacked transparents/particles/full-screen effects re-shading pixels.
No budgets. Without per-subsystem ms and asset ceilings, performance silently regresses;
enforce them in your build/CI checks.
Optimizing too early. Don't contort a prototype for performance before it's fun or measured.
References
For per-engine profiler walkthroughs, the CPU-vs-GPU triage flowchart, a complete pooling
manager, batching/instancing rules per engine, allocation/GC guidance, LOD/culling, and asset
budgets (texture sizes, triangle counts, audio, mobile thermals), read
references/profiling-and-budgets.md.
Related skills
physics-tuning — simulation cost, fixed-step budget, sleeping bodies, broadphase layers.
godot-export — release/build settings that affect measured performance.
procedural-gen, game-ai — common CPU hotspots (generation, pathfinding) to budget and defer.
roguelike, tower-defense, survival-crafting — entity-heavy genres that need pooling/budgets.don't have the plugin yet? install it then click "run inline in claude" again.
performance optimization is a measurement-driven discipline. measure first, identify the actual bottleneck (not the assumed one), fix it, then re-measure to confirm the fix worked. only optimize what data proves matters. premature optimization adds complexity without improving user experience. use this skill when performance requirements exist in the spec, when users report slow behavior, when core web vitals fall below thresholds, or when a change introduced a regression.
measurement tools (choose based on use case)
performance budgets (thresholds)
core web vitals baselines
external connections
baseline context
identify what is slow using the symptom tree below. match your user complaint (slow page load, sluggish interaction, etc.) to a category.
choose measurement method based on symptom type:
npx lighthouse <url> --output=json)console.time('db-query') and console.timeEnd('db-query'))import { onLCP, onINP, onCLS } from 'web-vitals';
onLCP(sendToAnalytics);
onINP(sendToAnalytics);
onCLS(sendToAnalytics);
run the measurement in fixed conditions:
output: baseline metrics document with timestamp, conditions, and raw numbers (not smoothed or cherry-picked).
symptom tree (where to start measuring)
what is slow?
├── first page load
│ ├── large bundle? → measure bundle size, check code splitting
│ ├── slow server response? → measure TTFB in devtools network waterfall
│ │ ├── DNS long? → add dns-prefetch / preconnect for known origins
│ │ ├── TCP/TLS long? → enable HTTP/2, check edge deployment, keep-alive
│ │ └── waiting (server) long? → profile backend, check queries and caching
│ └── render-blocking resources? → check network waterfall for CSS/JS blocking
├── interaction feels sluggish
│ ├── UI freezes on click? → profile main thread, look for long tasks (over 50ms)
│ ├── form input lag? → check re-renders, controlled component overhead
│ └── animation jank? → check layout thrashing, forced reflows
├── page after navigation
│ ├── data loading? → measure API response times, check for request waterfalls
│ └── client rendering? → profile component render time, check for N+1 fetches
└── backend / API
├── single endpoint slow? → profile database queries, check indexes
├── all endpoints slow? → check connection pool, memory, CPU
└── intermittent slowness? → check for lock contention, GC pauses, external deps
map your symptom to the likely cause table below.
gather profiling evidence:
node --prof app.js then node --prof-process isolate-*.log), or use an APM tool to attribute latency to functionsEXPLAIN ANALYZE in postgres, similar in mysql/mongo), look for full table scans, missing indexes, or N+1 patternsnarrow to one specific bottleneck. do not attempt to fix multiple issues in one change; each fix must be measurable in isolation.
output: bottleneck identification with evidence (e.g., "database query taking 800ms, full table scan on users table, missing index on user_id").
bottleneck reference
frontend symptoms:
backend symptoms:
implement a fix targeting only the bottleneck identified in step 2. use the anti-pattern reference below as a guide. commit the change with a clear message describing what is being fixed and why (e.g., "add include on user relation to fix N+1 in task list endpoint").
common anti-patterns and fixes
N+1 queries (backend)
// bad: one query per task for the owner
const tasks = await db.tasks.findMany();
for (const task of tasks) {
task.owner = await db.users.findUnique({ where: { id: task.ownerId } });
}
// good: single query with join/include
const tasks = await db.tasks.findMany({
include: { owner: true },
});
unbounded data fetching
// bad: fetching all records
const allTasks = await db.tasks.findMany();
// good: paginated with limits
const tasks = await db.tasks.findMany({
take: 20,
skip: (page - 1) * 20,
orderBy: { createdAt: 'desc' },
});
missing image optimization (frontend)
<!-- bad: no dimensions, no format optimization, no priority hint -->
<img src="/hero.jpg" />
<!-- good: hero/LCP image with art direction and resolution switching -->
<picture>
<!-- mobile: portrait crop (8:10) -->
<source
media="(max-width: 767px)"
srcset="/hero-mobile-400.avif 400w, /hero-mobile-800.avif 800w"
sizes="100vw"
width="800"
height="1000"
type="image/avif"
/>
<source
media="(max-width: 767px)"
srcset="/hero-mobile-400.webp 400w, /hero-mobile-800.webp 800w"
sizes="100vw"
width="800"
height="1000"
type="image/webp"
/>
<!-- desktop: landscape crop (2:1) -->
<source
srcset="/hero-800.avif 800w, /hero-1200.avif 1200w, /hero-1600.avif 1600w"
sizes="(max-width: 1200px) 100vw, 1200px"
width="1200"
height="600"
type="image/avif"
/>
<source
srcset="/hero-800.webp 800w, /hero-1200.webp 1200w, /hero-1600.webp 1600w"
sizes="(max-width: 1200px) 100vw, 1200px"
width="1200"
height="600"
type="image/webp"
/>
<img
src="/hero-desktop.jpg"
width="1200"
height="600"
fetchpriority="high"
alt="Hero image description"
/>
</picture>
<!-- good: below-the-fold image with lazy loading and async decoding -->
<img
src="/content.webp"
width="800"
height="400"
loading="lazy"
decoding="async"
alt="Content image description"
/>
unnecessary re-renders (react)
// bad: creates new object on every render, causing children to re-render
function TaskList() {
return <TaskFilters options={{ sortBy: 'date', order: 'desc' }} />;
}
// good: stable reference
const DEFAULT_OPTIONS = { sortBy: 'date', order: 'desc' } as const;
function TaskList() {
return <TaskFilters options={DEFAULT_OPTIONS} />;
}
// use react.memo for expensive components
const TaskItem = React.memo(function TaskItem({ task }: Props) {
return <div>{/* expensive render */}</div>;
});
// use useMemo for expensive computations (only when profiler shows re-compute is slow)
function TaskStats({ tasks }: Props) {
const stats = useMemo(() => calculateStats(tasks), [tasks]);
return <div>{stats.completed} / {stats.total}</div>;
}
large bundle size
// good: dynamic import for heavy, rarely-used features
const ChartLibrary = lazy(() => import('./ChartLibrary'));
// good: route-level code splitting wrapped in Suspense
const SettingsPage = lazy(() => import('./pages/Settings'));
function App() {
return (
<Suspense fallback={<Spinner />}>
<SettingsPage />
</Suspense>
);
}
missing caching (backend)
// cache frequently-read, rarely-changed data
const CACHE_TTL = 5 * 60 * 1000; // 5 minutes
let cachedConfig: AppConfig | null = null;
let cacheExpiry = 0;
async function getAppConfig(): Promise<AppConfig> {
if (cachedConfig && Date.now() < cacheExpiry) {
return cachedConfig;
}
cachedConfig = await db.config.findFirst();
cacheExpiry = Date.now() + CACHE_TTL;
return cachedConfig;
}
// HTTP caching headers for static assets
app.use('/static', express.static('public', {
maxAge: '1y', // Cache for 1 year
immutable: true, // Never revalidate (use content hashing in filenames)
}));
// cache-control for API responses
res.set('Cache-Control', 'public, max-age=300'); // 5 minutes
re-run the exact measurement from step 1 under identical conditions (same device, network, cache state, minimum 3 runs).
compare the new result against the baseline. do not average across old and new runs; keep them separate.
check if the improvement exceeds run-to-run variance. if variance is ±15ms and the improvement is 10ms, that is inside noise and does not count.
verify that all tests still pass (do not change tests to hide a regression).
output: before/after comparison with raw numbers, variance range, and decision (keep or revert).
apply this matrix strictly:
| result vs baseline | verdict |
|---|---|
| past the threshold AND tests green | keep. commit with before/after numbers in the message. |
| within noise (no measurable change) | revert. neutral changes are a revert. |
| worse | revert. |
| improved BUT a test went red | revert. a regression wearing a win's clothing is still a regression. |
sunk cost is not a reason to keep a change. if you wrote it and it did not work, revert it. the codebase accretes complexity that never bought anything, and you maintain it forever.
if the optimization was kept, add a performance assertion in CI:
npx bundlesize --config bundlesize.config.jsonnpx lhci autorun --config=lighthouserc.jsonlog the optimization attempt (keep and reverted alike) in a PERF.md or PR checklist so the same dead idea does not get re-run:
| idea | baseline → result | verdict | why |
|---|---|---|---|
| memoize row component | INP 240ms → 235ms | reverted | inside noise (±15ms). rows not the bottleneck. |
| virtualize list | INP 240ms → 90ms | kept | long tasks gone from trace. |
| preconnect to API origin | LCP 2.8s → 2.8s | reverted | already same-origin. |
when to optimize vs. when to skip
synthetic vs. RUM measurement
cold cache vs. warm cache baseline
change isolation
noise vs. signal
correctness gates metrics
memoization and useMemo red flags
step 1 (measure) deliverable
{ "metric": "LCP", "baseline_mean_ms": 2800, "baseline_min_ms": 2600, "baseline_max_ms": 3100, "conditions": "4G, cold cache, moto g4", "date": "2024-01-15" }**step 2 (