TechIDaily Journal
Real-Time Camera Inversion on Metal: Inside ReverseWorldGo
ReverseWorldGo flips your camera feed in real time using Metal compute shaders. We walk through the five-stage pipeline, the kernel we wrote, and the latency budget that made it feel native on every iPhone since the A12.
Real-Time Camera Inversion on Metal: Inside ReverseWorldGo's Vision Pipeline
*The five-stage GPU pipeline that flips your camera feed in under 16 milliseconds, on every iPhone since the A12.*
ReverseWorldGo is a puzzle game that asks players to see the world backwards. The camera is the controller, and inversion is the verb. What sounds like a one-line Core Image filter turned into a six-month engineering project because the moment you put a full-resolution live preview behind it, the frame budget collapses from sixty to nine frames per second. This post is the story of how we rebuilt the inversion pipeline on Metal, what it cost us in late nights, and the five mistakes that taught us the most.
1. Why Metal, Not Core Image
Our first implementation used CIFilter with a custom CIKernel and the framework's automatic Metal-backed renderer. It worked on paper. It also ran at nine frames per second on an iPhone 12 mini, and the thermal throttling curve meant that by minute three of a session the framerate dropped into single digits. The bottleneck was not the inversion itself — it was the implicit copy between the CVPixelBuffer from AVCaptureSession and the CIImage that Core Image wanted to render.
Metal removes that copy. A CVMetalTextureCache lets you wrap a CVPixelBuffer as a MTLTexture without copying the underlying pixel memory, and a compute kernel can read from that texture, write to a second texture, and present in a single GPU dispatch. The latency budget drops from a frame-and-a-half to a sub-frame, and the power envelope stays inside the thermal headroom of every A-series chip we support.
2. The Five-Stage Inversion Pipeline
The pipeline is intentionally flat. Each stage is a Metal object with a single responsibility, and the only shared state is the texture cache and the command queue.
Stage 1: AVCaptureSession → CVPixelBuffer (YCbCr 4:2:2)
Stage 2: CVMetalTextureCache → MTLTexture (.bgra8Unorm)
Stage 3: InversionComputeKernel → MTLTexture (inverted)
Stage 4: SpriteKit compose → MTLTexture (with UI overlay)
Stage 5: MTKView present → CAMetalLayer drawableEach stage runs in 1-3 milliseconds on an A15, and the pipeline as a whole finishes inside a single 16.67 ms display tick. There is no readback to the CPU, no staging buffer, and no autoreleasepool ping-pong.
3. Stage 1: Capturing the Camera
The capture session is configured for the highest preview resolution the device supports without breaking 60 fps. On a Pro device we run at 1920×1440. On the SE we drop to 1280×960. The session outputs kCVPixelFormatType_420YpCbCr8BiPlanarFullRange, which is the format the camera sensor delivers natively.
let session = AVCaptureSession()
session.sessionPreset = .high
let output = AVCaptureVideoDataOutput()
output.videoSettings = [
kCVPixelBufferPixelFormatTypeKey as String: kCVPixelFormatType_420YpCbCr8BiPlanarFullRange
]
output.alwaysDiscardsLateVideoFrames = true
output.setSampleBufferDelegate(self, queue: captureQueue)
session.addOutput(output)The alwaysDiscardsLateVideoFrames flag is the difference between a smooth preview and a stuttering one. Without it, the system queues buffers when our render loop falls behind, and the latency grows until the preview is two seconds behind reality.
4. Stage 2: Wrapping the Buffer as a Metal Texture
This is the trick that took the longest to find. CVMetalTextureCache is a system-managed pool that hands you a MTLTexture whose underlying memory is the CVPixelBuffer's memory. There is no copy, no staging, no allocation.
var cvTexture: CVMetalTexture?
CVMetalTextureCacheCreateTextureFromImage(
kCFAllocatorDefault,
textureCache,
pixelBuffer,
nil,
.bgra8Unorm,
CVPixelBufferGetWidth(pixelBuffer),
CVPixelBufferGetHeight(pixelBuffer),
0,
&cvTexture
)
let texture = CVMetalTextureGetTexture(cvTexture!)!The YCbCr buffer does need a YUV-to-RGB conversion before it reaches this stage, and we use a separate Metal kernel for that. The combined cost of YUV conversion and inversion is one dispatch.
5. Stage 3: The Inversion Kernel
The actual inversion is a single compute kernel written in the Metal Shading Language. It runs on every thread that maps to a pixel and produces the inverted color in the destination texture.
#include <metal_stdlib>
using namespace metal;
kernel void invertKernel(
texture2d<float, access::read> inTex [[texture(0)]],
texture2d<float, access::write> outTex [[texture(1)]],
uint2 gid [[thread_position_in_grid]]
) {
float4 color = inTex.read(gid);
color.rgb = 1.0 - color.rgb;
outTex.write(color, gid);
}We tried several variants — fragment shaders, render-to-texture pipelines, MPSImageColorInvert — and the custom compute kernel won on every benchmark. The fragment shader variant required a render pass descriptor and a vertex buffer, which added two milliseconds. MPSImageColorInvert was slightly faster but forced us into a different texture format that broke our downstream composition stage. The custom kernel is fast, simple, and trivially auditable.
6. Stage 4: Composition with the Game UI
ReverseWorldGo is not just an inverted camera. The puzzle layer is drawn on top, and the inversion applies only to the camera region. We use a MTKView for the camera and a SKView for the puzzle overlay, composited through a shared CAMetalLayer.
The trick is that SpriteKit does not know it is rendering into a Metal-backed layer that is also being read by another Metal pass. We had to set skView.preferredFramesPerSecond = 60 and ensure the overlay renders after the camera pass. The first version of this code rendered the overlay first, then the camera, and the result was a flickering UI that no amount of isHidden could fix.
7. Stage 5: Present to Display
The final stage is the simplest. We acquire the next drawable from MTKView.currentDrawable, schedule the command buffer, and present. The command buffer is the only synchronization point, and Metal's automatic dependency tracking handles the rest.
The latency from camera shutter to display is 2-3 frames, which is the floor for any camera-plus-render pipeline. Anything below 2 frames requires a specialized sensor path that the public AVCaptureSession API does not expose.
8. Battery and Thermal
The full pipeline draws 6-8% battery per hour of continuous play on an iPhone 14 Pro. The thermal curve plateaus at 38°C after fifteen minutes and stays there for the rest of the session. We confirmed this with three days of ProcessInfo.thermalStateDidChangeNotification logging across a fleet of test devices.
The single biggest thermal win was dropping the capture resolution from 4K to 1920×1440. The 4K preview looked marginally better but ran the GPU 22% hotter, and most players could not tell the difference in a side-by-side test. The lesson is that pixel count is rarely the right knob to turn when the bottleneck is thermal, not pixel-bound work.
9. Five Mistakes We Made
We shipped several decisions that did not work, and we want to be specific.
Render-to-texture first. Our first version rendered the camera into an offscreen texture, then sampled that texture to invert it. The intermediate readback cost 4 ms per frame and pushed us below 60 fps on the iPhone 12 mini. Compute kernels write directly to the destination texture and skip the readback.
Core Image "automatic" optimization. Core Image's automatic Metal backend does not always pick the optimal path. Our custom CIFilter ran faster as a Core Image filter than as a Metal compute kernel for some operations, but the auto-detection got it wrong on two of three devices we tested. We moved to Metal explicitly and stopped trusting the auto-detection.
YCbCr to BGRA in the kernel. The first version of our YUV-to-RGB conversion ran in the same kernel as the inversion. It looked elegant but the threadgroup memory pressure broke on the A12, which has a smaller L1 cache. We split the conversion into its own kernel and the A12 went from 14 fps back to 60 fps.
AVCaptureSession on the main thread. The session start method blocks for 100-300 ms while the sensor initializes. We initially called it from viewDidAppear and the first frame stuttered. Moving the start to viewDidLoad, before the user can see the screen, removed the stutter entirely.
SpriteKit overlay last. We rendered the puzzle overlay first, then the camera on top. The camera covered most of the overlay and the puzzle elements flickered as they were partially occluded. Rendering the camera first and the overlay on top fixed it.
10. What's Next
Three things are in active development. First, a visionOS variant that anchors the inverted feed to a fixed plane in the user's space. Second, a portrait-mode depth variant that inverts foreground and background independently for an occlusion-aware puzzle mode. Third, a multiplayer mode where two players on the same Wi-Fi network share an inverted feed and have to coordinate a solution.
If any of this resonates with how you build GPU pipelines on iOS, ReverseWorldGo is free to download with an optional Pro tier that unlocks the multiplayer mode. The free tier ships with twenty puzzles and the full inversion pipeline, so you can see the architecture in action without spending a cent.