Converting my Heaphones to work with Bluetooth

From a Burning Ear Cup to Hi-Fi Wireless: Building the BlueAudio Puck on an ESP32
Embedded Audio & Hardware Mod • Bluetooth Classic BR/EDR • ESP-IDF v5.5 • PCM5102A I²S DAC

From a Burning Ear Cup to Hi-Fi Wireless: Building the BlueAudio Puck on an ESP32

How a swollen battery replacement disaster on an intern keepsake led to building a pocket-sized, DSP-powered Bluetooth A2DP audio receiver using an original ESP32, an audiophile I²S DAC, and zero dongles.

#01 The Cherished Intern Keepsake Meets a Fiery End

During my engineering internship at Vega Innovations, I received a humble set of wireless headphones: an Airphone 3. While inexpensive, it was surprisingly robust. For nearly three and a half years of daily engineering work, PCB troubleshooting, and commuting, it served me faithfully across both Bluetooth and its 3.5 mm analog input without skipping a beat.

Eventually, Father Time caught up with the chemistry. After three years of constant charge-discharge cycles, the internal lithium cell stopped holding charge and puffed up into a swollen pillow.

Airphone 3 Wireless Headphones
The trusty Airphone 3: 3.5 years of dependable audio
Opening the left earcup assembly
Disassembling the left earcup containing the battery and receiver PCB

Thinking it would be an easy weekend repair, I sourced a replacement 3.7V 500mAh LiPo pouch cell from Tronic.lk and soldered it into the earcup power rails.

3.7V 500mAh Replacement LiPo Cell
The replacement 3.7V 500mAh LiPo cell from Tronic.lk
The Heat Meltdown: Almost immediately after powering on with the new cell, the factory Bluetooth SoC on the headphone circuit board began heating up uncontrollably. Within minutes, the plastic shell reached thermal runaway temperatures—it was so fiercely hot against my skull that my left ear was literally sweating!

Whether due to an internal silicone breakdown in the PMIC or an over-voltage quirk on the budget Bluetooth silicon, the integrated radio board was a legitimate safety and fire hazard. I desoldered the battery permanently and relegated the Airphone to being a dumb, passive wired headset over its 3.5 mm auxiliary jack.

#02 The Modern Smartphone Dilemma: Dongle Hell

Reverting to an analog wire introduced a frustrating modern problem: smartphones no longer have 3.5 mm headphone jacks.

Carrying fragile USB-C dongles everywhere is infuriating. They strain your phone’s charging port, get forgotten on desks, and physically tether your head to your workstation whenever you walk around the lab or fetch coffee.

The drivers inside the Airphone 3 were still in pristine mechanical shape. Rather than throwing out great acoustic hardware, I decided to engineer a dedicated, pocketable wireless bridge: The BlueAudio Puck.

#03 The Silicon Trap: Why the Original ESP32 Beats Modern Chips

When planning a modern Bluetooth audio receiver, most engineers instinctively grab Espressif’s newer microcontrollers—the ESP32-S3, C3, C5, or C6. They feature newer RISC-V or modern Xtensa cores, USB native OTG, and lower power consumption.

However, if you try building an A2DP Bluetooth audio receiver on those newer chips, you will hit a brick wall.

⚡ Why A2DP Sink Requires the Classic Dual-Core ESP32 ▼

Standard Bluetooth audio streaming from smartphones and computers relies on the Advanced Audio Distribution Profile (A2DP). A2DP runs exclusively on Bluetooth Classic (BR/EDR) physical and link layers.

  • Espressif dropped Bluetooth Classic completely on the ESP32-S3, C3, and C6 to save silicon die space and focus on IoT mesh, making them BLE-only.
  • They cannot fall back to the newer Bluetooth LE Audio standard either. LE Audio requires isochronous channels specified in Bluetooth 5.2. The S3 and C3 are locked to Bluetooth 5.0 hardware architectures.
  • Therefore, the venerable, original dual-core ESP32 (Xtensa LX6) remains the reigning king for DIY A2DP audio streaming.
Chip Series BT Radios A2DP Sink Capability Audio Verdict
ESP32 (Original) Classic BR/EDR + BLE 4.2 ✅ Full Native A2DP Sink & Source Ideal candidate for custom audio pucks
ESP32-S3 BLE 5.0 only ❌ Not supported (No BR/EDR) Cannot pair as a Bluetooth headphone
ESP32-C3 / C6 BLE 5.0 / 5.3 ❌ Not supported (No BR/EDR) Cannot sink standard mobile A2DP

#04 The Hardware Architecture: PCM5102A & Direct-Solder Sandwich

While the ESP32 has internal 8-bit DACs on GPIO 25 and 26, their dynamic range and signal-to-noise ratio (SNR) are abysmal for high-fidelity audio.

To get clean, audiophile-grade output, I paired the ESP32 with an external Texas Instruments / Burr-Brown PCM5102A 32-bit / 384 kHz I²S DAC. Rather than messy jumper wires, the pin mapping was intentionally routed so the PCM5102A module sits directly beneath the ESP32 devkit on rigid solder bridges.

ESP32 Pin Bus Signal Target Destination Function & Hardware Notes
GPIO 4 I²S BCK PCM5102A BCK Continuous bit clock line
GPIO 15 I²S LRCK PCM5102A LRCK Word select / Left-Right frame clock
GPIO 2 I²S DATA PCM5102A DIN Serial PCM audio stream (also pulses on-board LED!)
GPIO 21 / 22 I²C SDA / SCL SSD1306 OLED 128x64 display telemetry bus (0x3C address)
GPIO 33 RTC Button (BT2) Play / Pause / Pair Wakes the ESP32 from deep sleep on press

⚠️ PCM5102A Solder Jumpers: Solving the "No Audio" Trap

The PCM5102A requires specific hardware strap jumpers on its reverse side:

  • SCK Bridge Soldered: Tells the DAC’s internal PLL to generate its own master clock (MCLK) directly from BCK. This completely eliminates the need for an external MCLK line!
  • XSMT tied to 3.3V: Releases the soft-mute circuit. (Floating or pulled low results in complete silence).
  • FMT tied to GND: Selects standard I²S Philips data format.
  • GPIO 2 Strapping Note: Because GPIO 2 is a boot strapping pin, holding the board's BOOT button may be required when flashing new firmware over UART.
BlueAudio Puck ESP32 Prototype
The BlueAudio Puck: Compact ESP32 + PCM5102A receiver driving the Airphone 3

#05 Firmware Architecture: Real-Time Audio Without Stutters

Streaming audio over a microcontroller running an operating system is a real-time juggling act. If the Bluetooth radio gets starved for even a couple of milliseconds, packets drop and audio stutters.

Built from scratch on ESP-IDF v5.5.4, the firmware adheres to a strict architectural rule: Bluedroid stack callbacks are non-blocking and only enqueue.

Firmware Audio Pipeline Execution Pinned Dual-Core Model
[Bluetooth Controller / Radio Stack] ➔ Core 0
                 │ (Raw SBC Packets over ACL)
                 ▼
[Bluedroid Decoder Callback] (Instant Queue Push)
                 │
                 ▼ (FreeRTOS 32 kB Ring Buffer with 25% Prefetch Cushion)
                 │
[Audio Writer & DSP Task] ➔ Pinned to Core 1
                 │
                 ├─ 5-Band Biquad Parametric EQ Cascade (1337 µs / 360 frames)
                 ├─ Software Volume Scaling (AVRCP Synchronized)
                 ▼
[I²S DMA Controller (APLL Clocked @ 44.1 kHz)] ➔ PCM5102A DAC ➔ 3.5mm Jack

By pinning the audio writer, DSP filters, and blocking I²S calls to Core 1, the radio on Core 0 is never interrupted. Furthermore, using the ESP32’s internal Audio PLL (APLL) generates clean fractional clock dividers, producing an exact 44.1 kHz sample rate with virtually zero clock jitter.

🎛️ Integrated 5-Band Parametric Equalizer

Because budget headphone drivers often have resonant peaks or rolled-off sub-bass, the firmware includes a real-time 5-band biquad parametric equalizer.

Benchmarked on hardware, calculating all 5 stereo biquad sections takes just 1,337 µs per 360-frame audio block—consuming only ~16% of Core 1 at 160 MHz. This leaves ample CPU headroom while allowing complete acoustic correction directly in firmware!

#06 Display Telemetry, AVRCP & Paying It Forward

To make the Puck feel like a polished commercial consumer device, I integrated a 0.96-inch SSD1306 128x64 OLED display communicating over I²C:

  • Full AVRCP Metadata: Displays current track title, artist name, and playback transport state (Play/Pause). Overlong song titles automatically bounce and scroll smoothly.
  • Absolute Volume Synchronization: Changing the volume on your smartphone updates the Puck's on-screen volume bar in real-time, and vice-versa.
  • Golden Receive Power Range (RSSI): Rather than meaningless fake dBm bars, the link meter reads Bluetooth Classic's delta from the optimal RF reception window.
  • Power Management: Dynamic Frequency Scaling switches the CPU between 80 MHz and 160 MHz, transitioning into deep sleep after 15 minutes of inactivity.

Seeing how convenient and clear the audio was, I built a secondary custom unit for a friend who was dealing with the exact same missing-jack predicament on his setup:

Secondary Bluetooth Receiver Build
Dedicated standalone receiver built for a friend

#07 Conclusion & Open Source Repository

What started as an alarming near-fire on an ear cup ended with an audio receiver that delivers significantly higher audio quality and lower noise than the factory Airphone 3 ever possessed.

By leveraging dedicated hardware DACs, APLL audio clocks, and smart FreeRTOS core distribution, you can breathe new, wireless life into any vintage or favorite wired headphones without sacrificing sound quality or buying throwaway plastic dongles.

Get the Firmware & KiCad Schematics

The complete ESP-IDF v5.5 project, parametric DSP code, and KiCad PCB designs are available on GitHub.

View Repository →

Published with ❤️ by the hardware engineering community • Powered by ESP32 & FreeRTOS

Making Most Out of my Old AMD Laptop

From Broken Hinge to 35W Homelab Server: Resurrecting a Huawei MateBook 13
Hardware Mod & Embedded Linux • AMD Zen+ Picasso • ACPI DSDT • 24/7 Compute

From Broken Hinge to 35W Homelab Server: Resurrecting a Huawei MateBook 13

How reverse-engineering AMD’s Precision Boost 2, ACPI tables, and the Embedded Controller turned an abandoned ultrabook with dead cells into an overclocked, headless compute powerhouse.

#01 The College Workhorse Meets Its Demise

Back in 2021, I bought a truly solid laptop: the Huawei MateBook 13 AMD. Powered by an AMD Ryzen 5 3500U (4 cores, 8 threads), 16 GB of dual-channel DDR4 RAM, and a 512 GB NVMe SSD, it was a sleek, lightweight machine that served as my primary engineering workstation throughout university.

From multi-layer PCB design and embedded electronics programming to compiling software and gaming, it handled heavy workloads without breaking a sweat. Its integrated Radeon Vega 8 graphics were surprisingly capable, comfortably pushing titles like GTA V at 720p 60 FPS.

It felt like an aluminum MacBook Air clone, but built with genuine thermal headroom: beefy power VRMs and a dual-fan cooling assembly allowed the CPU to boost up to its 3.5 GHz single-core limit over its 2.1 GHz base clock. But after four years of daily transport, the chassis hinge snapped clean off, and shortly thereafter, the internal battery pack died completely.

Removing broken lid
Power house
Motherboard inspection
University
Bench testing
Laptop Brand new from Store

#02 Turning E-Waste into a 24/7 Compute Server

Rather than discarding capable quad-core x86 silicon, I stripped away the broken display assembly and turned the board into a 24/7 headless Linux server running Ubuntu.

To test its endurance, I assigned it three continuous background workloads:

I2P Network Logo

I2P Router

Decentralized network routing node helping preserve privacy and connectivity.

ArchiveTeam Warrior Logo

ArchiveTeam Warrior

Distributed web crawling workers archiving at-risk digital history.

BOINC Distributed Compute Logo

BOINC Distributed

Multi-core AVX PrimeGrid sieves and Radeon Vega 8 OpenCL mathematical transforms.

#03 Hitting the Dreaded 400 MHz Throttling Wall

Everything looked great until sustained loads began. That is when I ran into a massive firmware restriction: without a functional battery attached, the motherboard refused to boost past 2.1 GHz. The GPU was pegged at 450 MHz, and periodically the entire system throttled violently down to 400 MHz (0.4 GHz).

The Electrical Cause: Modern laptops use the battery as a high-current surge buffer. When the Embedded Controller (EC) cannot communicate with a healthy battery, it trips emergency safety downclocking via a hardware line named BD PROCHOT# (Bi-Directional Processor Hot) to prevent power adapter brownouts.

My initial instinct was straightforward: rebuild the pack by wiring new high-capacity 18650 lithium cells to the factory battery management board.

Rebuilt 18650 battery pack mod
Wiring rebuilt 3S 18650 cells to the factory Huawei BMS circuitry

However, the system still fell into 400 MHz to 1.5 GHz drops. This is where hardware meets firmware safety: laptop Battery Management Systems (BMS) store state flags in volatile memory. If cells drop below critical voltage or disconnect, the BMS locks out irreversibly to prevent fires. Replacing the cells had triggered a permanent limp mode—no charging, no discharging, and continuous alert signals sent to the motherboard!

#04 Reverse Engineering the ACPI Tables & EC RAM

Since the physical battery circuit had locked itself, I disconnected it completely and attacked the problem at the firmware interface.

By dumping and disassembling the ACPI DSDT table (/sys/firmware/acpi/tables/DSDT), I found exactly how the operating system and the Embedded Controller evaluate battery status:

🔍 Deep Technical Breakdown: The ACPI _BIX Check & EC RAM Map ▼

In the disassembled DSDT, the ACPI Extended Battery Information method (_BIX) contains a strict evaluation condition:

Method (_BIX, 0, Serialized) {
    If (ECOK ()) {
        If ((Acquire (Z009, 0x2000) == Zero)) {
            // CRITICAL SANITY CHECK: ALL THREE MUST BE NON-ZERO
            If (((BTDV && BTFC) && BTDC)) {
                BPKG [0x02] = BTDC /* Design Capacity (3610 mAh) */
                BPKG [0x03] = BTFC /* Last Full Charge Capacity */
                BPKG [0x05] = BTDV /* Design Voltage (11400 mV) */
            }
            Release (Z009)
        }
    }
    Return (BPKG)
}

Without a battery, the EC RAM offsets for BTDC (0x84) and BTDV (0x86) remain at 0x00. This causes the kernel driver to omit capacity from sysfs, prompting system monitors like btop to disable CPU boost scaling.

EC Offset Field Injected Value Function
0x80 ACST / BST1 0x03 Signals AC Online + Battery Present
0x84 BTDC 3610 mAh Design Capacity (Unlocks _BIX)
0x86 BTDV 11400 mV Design Voltage (11.4V nominal)
0x90 BAPV 12500 mV Present Pack Voltage (12.5V simulated)

I developed a lightweight background service, ec_voltage_patcher.py, that writes these register values directly into the Embedded Controller's RAM using the Linux kernel's ec_sys interface at 50 Hz (20ms intervals).

⚡ Calibrating the ACPI IRQ 9 Loop

Writing to the EC triggers hardware interrupts over LPC/eSPI. Running a tight 5ms loop initially hammered 3,200 interrupts/sec, burning 20% of an entire CPU core on [irq/9-acpi]! Calibrating the interval to 20ms (50 Hz) dropped interrupts by 75%, reclaiming ~15% of that core for compute while keeping telemetry smooth.

#05 Unlocking 35W Power & Eliminating 400 MHz Drops

By default, Huawei firmware chokes the Ryzen 5 3500U to an 18W STAPM limit. With battery emulation active, I leveraged ryzenadj to command the processor's System Management Unit (SMU) registers directly:

SMU Power Tuning Command Active in daemon
ryzenadj \
  --stapm-limit=35000 \        # 35W Sustained Power Target
  --fast-limit=35000 \         # 35W Fast PPT (Matched to prevent VRM over-current)
  --slow-limit=35000 \         # 35W Slow PPT
  --vrm-current=55000 \        # 55A TDC Sustained Current Limit
  --vrmmax-current=70000 \     # 70A EDC Peak Electrical Current Limit
  --tctl-temp=88 \             # 88°C Thermal Ceiling
  --prochot-deassertion-ramp=1 # Collapses 400 MHz recovery from 30s to 1ms!

The key parameter here is --prochot-deassertion-ramp=1. Normally, when a transient spike trips the processor hot line, AMD's default firmware enforces a sluggish 30-second penalty box at 400 MHz. Setting the ramp to 1 forces the processor to recover in just 1 millisecond, preventing random clock stalls.

#06 Overcoming Heat-Soak: The PTM7950 Repaste

While power limits were unlocked, the hardware soon met another roadblock: thermal throttling. At 35W, the bare silicon die was shooting past 85°C almost instantly.

Disassembling the stock copper cooling assembly revealed the culprit: after 4 years of heavy use, the factory thermal paste had baked out completely into dry, chalky powder, creating micro-air gaps between the bare silicon and copper coldplate.

Dried thermal paste on heatsink
The factory thermal paste had dried into a chalky thermal insulator

To address this permanently, I completed three key hardware upgrades:

  • Honeywell PTM7950 Phase-Change Pad: Melting at ~45°C to fill microscopic voids on the bare silicon, it provides near liquid-metal heat transfer with zero pump-out effect and zero electrical short risks.
  • High-Conductivity VRM Thermal Pads: Applied to the power delivery MOSFETs to dissipate heat directly into an external aluminum heatsink with active fan cooling.
  • 100W USB-PD Power Supply: Upgraded from the factory 65W charger to a 100W brick, preventing transient voltage sag that previously tripped emergency BD PROCHOT downclocks.
Stable 35W All-Core Boost Telemetry
Stable 35W sustained operation: all 8 threads holding 2.7 to 3.1 GHz with temps at an ultra-cool 66°C

#07 Profile Comparison Matrix & Verified Results

Here is how the tuned profiles compare across different operating conditions:

Profile Fast PPT Slow PPT Tctl Temp All-Core Clocks Noise & Thermals Recommended For
Silent / Cool 25 W 22 W 72 °C ~2.40 – 2.50 GHz Whisper quiet, ~65–70°C Office work, quiet environments
Stabilized 65W 30 W 28 W 84 °C ~2.40 – 2.89 GHz Passive block / OEM brick 24/7 compute with OEM 65W adapter
Active Fan + 100W PD
Current Active
38 W 35 W 88 °C ~2.72 – 3.16 GHz Cool (~66–75°C) with fan 24/7 BOINC (CPU+GPU) + Frigate NVR
Max Turbo / Water 45 W 45 W 95 °C ~3.05 GHz (Physical limit) High airflow or loop, ~75–80°C Benchmarking & full custom liquid cooling
Sustained Package
35W
+133% over stock TDP
All-Core Clocks
~2.87 GHz
Up to 3.16 GHz burst
GPU VRAM Bus
1.2 GHz
38.4 GB/s pinned
Junction Temp
~66 °C
PTM7950 + Fan

#08 Conclusion: Giving E-Waste a High-Performance Second Life

What began as an e-waste pile of broken aluminum and dead battery cells evolved into an immensely rewarding journey through x86 power architecture and firmware reverse-engineering.

By combining EC RAM emulation, Precision Boost 2 tuning, and modern phase-change thermal materials, this old laptop now runs faster, cooler, and more reliably than it ever did fresh out of the factory box. Today, it quietly crunches distributed science calculations and manages live camera streams 24/7 without missing a beat.

Looking for the scripts and firmware?

The battery emulator daemon, systemd service, and ESP32 firmware are open-source on GitHub.

View Repository →

Published with ❤️ by the hardware engineering community • Powered by Linux & Open Source

Taming AMD's Linux AI Stack: From Kernel Panics to 80% Idle

Taming AMD's Linux AI Stack: From Kernel Panics to 80% Idle
Debugging • Linux Kernel • Computer Vision

Taming AMD's Linux AI Stack:
From Kernel Panics to 87% CPU Idle

How I fixed a fatal GPU deadlock in Frigate NVR by ripping out ROCm and routing AI inference through Linux gaming drivers.

The Short Version (For Everyone)

I recently set up a smart security camera system called Frigate on my home server. Frigate is incredibly smart—it looks at camera feeds in real-time to detect people, cars, and animals. To do this without melting the server's main processor (CPU), it uses the graphics card (GPU).

My server has a brand new AMD Ryzen processor with built-in graphics (Radeon 760M). On paper, it's a beast. In reality? The server was completely crashing and freezing every 6 to 10 minutes. I had to pull the power plug to fix it.

Why was it crashing?

Imagine a busy intersection with traffic lights controlled by a highly complex, proprietary computer system built by AMD (called ROCm). The system was trying to route two massive fleets of trucks at exactly the same time: one fleet carrying video data, the other carrying AI math calculations. The traffic controller completely panicked, caused a massive pileup, and then the tow trucks (the system reset protocol) broke down on the way to the scene.

How did I fix it?

Instead of relying on AMD's proprietary AI traffic controller, I fired them. I found a different, open-source tool (called ncnn) that routes the AI math through Vulkan. Vulkan is the exact same underlying technology that makes massive 3D video games run smoothly on Linux (like on the Steam Deck).

Because the gaming drivers are heavily tested by millions of players, they are rock solid. They handled the video and the AI math perfectly. My server went from crashing every 6 minutes and using 100% of its CPU, to running flawlessly with the CPU sitting at 87% idle, sipping power.


The Deep Dive (For the Engineers)

1. The Nightmare: D State and TTM Deadlocks

The hardware: An AMD Ryzen 5 8600G (Phoenix1 architecture, RDNA3, gfx1103 APU). The software: Dockerized Frigate 0.17 utilizing ONNX Runtime.

Frigate utilizes the GPU for two distinct pipelines: VAAPI for hardware video decoding (4 camera streams), and ROCm/MIGraphX for ONNX Runtime object detection (YOLOv9).

Shortly after startup, the Frigate container would become completely unresponsive. Docker could not kill it (SIGKILL was ignored). A quick dive into the host system revealed the horror:

USER       PID %CPU %MEM    VSZ   RSS TTY      STAT START   TIME COMMAND
root    337200  0.0  0.0      0     0 ?        D<   02:01   0:00 [kworker/u49:9+ttm]
root    337215  0.0  0.0      0     0 ?        D<   02:01   0:00 [kworker/u49:10+ttm]
... (16 workers stuck in D state)

Processes stuck in D (uninterruptible sleep) usually indicate severe I/O or kernel-level locks. Looking at dmesg confirmed it:

[37456.170181] amdgpu 0000:0f:00.0: GPU reset begin!. Source: 5

The Root Cause: Because this is an APU, system RAM is shared as VRAM. The VAAPI processes (media engine) and the MIGraphX processes (compute engine) were stepping on each other inside the TTM (Translation Table Maps) memory manager. This deadlock triggered a GPU reset (Source 5 = compute engine hang). However, on gfx1103, the amdgpu kernel reset sequence is known-buggy and never completed, resulting in a zombie GPU.

2. The Failed Attempts & ROCm Pitfalls

Before abandoning ROCm entirely, I tried standard mitigation strategies to stop the engines from fighting:

  • Attempt 1 (Software Decode + ROCm): I disabled VAAPI to dedicate the GPU strictly to MIGraphX inference. While ROCm achieved a blazing fast ~14ms inference speed, the GPU still hung after about 6 minutes. ROCm compute on consumer RDNA3 APUs on Linux is just fundamentally unstable, suffering from missing engine isolation.
  • Attempt 2 (VAAPI Decode + CPU Inference): I pushed YOLO detection to the CPU via standard ONNX Runtime and left VAAPI running on the GPU. Stability was achieved, but inference ballooned to ~77ms, and the CPU load hovered at a staggering 45 (94% to 100% utilization, completely pegged). Unacceptable for a home lab meant to run other services concurrently.

3. The Breakthrough: Vulkan and RADV

While AMD's proprietary compute stack (ROCm) is deeply flawed for consumer APUs, their Linux gaming stack is incredible. Mesa's RADV Vulkan driver is bulletproof. If I could route AI inference through Vulkan instead of ROCm, I could bypass the broken amdgpu compute paths entirely.

Microsoft's ONNX Runtime does have a Vulkan Execution Provider, but it is not shipped in their pre-built Linux wheels. Building it from source inside the Frigate container was an option, but a heavy one.

Instead, I pivoted to ncnn, Tencent's high-performance neural network inference framework optimized for mobile platforms. ncnn has native, highly-optimized Vulkan support. Because Python bindings for ncnn exist, I could write a custom detector plugin for Frigate.

4. Building the Custom Detector

I downloaded a pre-converted YOLOv5s .ncnn model and its parameter file. However, substituting ONNX for ncnn meant I lost Frigate's built-in ONNX post-processing. `ncnn` outputs raw logits; it doesn't apply sigmoid activations or grid decoding internally like the ONNX graphs do.

I wrote a custom Python class to hijack Frigate's ONNXDetector type, manually reshape the raw tensors, apply the sigmoid functions, map the anchor boxes, and run Non-Maximum Suppression (NMS). Here is a snippet of the crucial grid decoding logic that bridged the gap:

# Reshape YOLOv5 outputs: (255, H, W) -> (3, 85, H, W) -> (3*H*W, 85)
def decode_output(ncnn_mat, stride):
    arr = np.array(ncnn_mat)  # (255, grid_h, grid_w)
    na, nc = 3, 80
    no = 5 + nc  # 85 (4 box + 1 obj + 80 class)
    grid_h, grid_w = arr.shape[1], arr.shape[2]
    
    # Reshape and permute
    arr = arr.reshape(na, no, grid_h, grid_w)
    arr = np.transpose(arr, (2, 3, 0, 1))  # (grid_h, grid_w, 3, 85)
    arr = arr.reshape(-1, no)  # (grid_h*grid_w*3, 85)
    
    # Apply sigmoid (ncnn outputs raw logits)
    arr = 1.0 / (1.0 + np.exp(-arr))
    
    # Generate anchor grids
    grid_y, grid_x = np.meshgrid(np.arange(grid_h), np.arange(grid_w), indexing='ij')
    grid = np.stack([grid_x, grid_y], axis=-1)
    grid = np.expand_dims(grid, axis=2)
    grid = np.tile(grid, (1, 1, na, 1)).reshape(-1, 2)
    
    # Extract bounding boxes
    xy = arr[:, :2]
    wh = arr[:, 2:4]
    
    # YOLOv5 Decode formula
    xy = (xy * 2.0 - 0.5 + grid) * stride
    anchors_tiled = np.tile(self._anchors[stride], (grid_h * grid_w, 1))
    wh = (wh * 2.0) ** 2 * anchors_tiled
    
    # ... Confidence thresholding and NMS follows ...

5. Final Results: ROCm vs CPU vs Vulkan

With the custom image deployed, the transformation was instantaneous. By running the inference directly through Mesa's RADV gaming drivers, we kept the performance benefits of hardware acceleration while dodging the kernel panics entirely.

Metric ROCm + VAAPI
(The Original Goal)
CPU + VAAPI
(The Fallback)
ncnn Vulkan + VAAPI
(The Fix)
Inference Speed ~14ms 🚀 ~77ms 🐌 ~28ms ⚡
CPU Idle ~74% 0% (Completely Pegged) 87%
GPU Status Deadlock / Hang Video decode only Decode + Inference (~39% Util)
System Stability ❌ Kernel Panic (6 min) ✅ Stable (But system unusable) ✅ 100% Stable (3+ hrs)

The Legacy Hardware Implication: This workaround isn't just for bleeding-edge RDNA3. Older APUs like the Ryzen 3500U (Vega 8) were never officially supported by ROCm. However, because they fully support Vulkan 1.2, this exact same ncnn Vulkan pipeline enables hardware-accelerated AI on them flawlessly.

The PR for this fallback mechanism is open on GitHub. If you're running Frigate on AMD hardware and pulling your hair out over amdgpu kernel crashes, bypass ROCm entirely. Vulkan is the way.

© 2026 Engineering Log. Building resilient systems on Linux.

See Through Walls with a $9 Microcontroller

See Through Walls with a $9 Microcontroller | RuView
router
cell_wifi Edge AI & Sensors

See Through Walls with a $9 Microcontroller

Deploying WiFi DensePose on Kubernetes.

"How I deployed a real-time WiFi-based human sensing system on a homelab K3s cluster with a $9 ESP32-S3, live pose estimation, and RTSP camera fusion."

WiFi signals pass through walls. When a person moves — or even breathes — those signals scatter differently. What if you could read that scattering pattern and reconstruct what happened on the other side?

That's exactly what RuView does. Built on research from Carnegie Mellon's DensePose From WiFi paper, RuView is an open-source edge AI system that turns commodity WiFi signals into real-time human pose estimation, vital sign monitoring, and presence detection — all without a single pixel of video.

I took it a step further: deployed it on Kubernetes, wired up a live ESP32-S3 sensor, and fused the WiFi signal data with an RTSP camera feed for dual-modal pose estimation. Here's how.

memory

The Hardware: $9 and a WiFi Router

The entire sensing hardware cost me $9:

  • check_circle 1x ESP32-S3 ($9) — a dual-core microcontroller with WiFi that exposes Channel State Information (CSI). CSI gives you per-subcarrier amplitude and phase data — 56+ data points per WiFi frame, 20 times per second. That's the raw material for sensing.

Standard consumer WiFi only gives you RSSI (a single signal strength number). CSI is like going from a thermometer to a thermal camera — instead of one number, you get a detailed map of how the signal is being affected by everything in the room.

info

I also had three ESP32-C3s sitting around, but those are single-core RISC-V chips that can't handle the DSP pipeline. The S3's dual-core Xtensa is required — one core captures CSI interrupts while the other runs signal processing.

layers

The Software Stack

RuView's Rust sensing server processes the signal chain:

terminal
ESP32 CSI (UDP) → Hampel outlier rejection → SpotFi phase correction
    → Fresnel zone modeling → FFT vital sign extraction
    → AI backbone (RuVector attention networks)
    → 17 body keypoints + breathing rate + heart rate + presence

At 54,000 frames/sec throughput in Rust, this is fast enough to process live data from multiple sensors with headroom to spare. The server exposes a REST API, WebSocket stream, and a full browser UI.

bolt

Flashing the ESP32-S3

The firmware ships as pre-built binaries in GitHub Releases. Flashing takes 30 seconds:

Bash
pip install esptool
python -m esptool --chip esp32s3 --port COM7 --baud 460800 \
  write_flash --flash_mode dio --flash_size 8MB \
  0x0 bootloader.bin \
  0x8000 partition-table.bin \
  0xd000 ota_data_initial.bin \
  0x10000 esp32-csi-node.bin

Then provision it with your WiFi credentials and the IP of your server:

Bash
python provision.py --port COM7 \
  --ssid "MyWiFi" --password "secret" \
  --target-ip 192.168.1.9

The ESP32 connects to WiFi and starts streaming CSI frames over UDP to port 5005. No internet needed after provisioning — everything stays local.

view_in_ar

Containerizing for Kubernetes

The project includes a multi-stage Dockerfile that compiles the Rust server and bundles the UI into a minimal Debian image:

Dockerfile
FROM rust:1.85-bookworm AS builder
WORKDIR /build
COPY rust-port/wifi-densepose-rs/ ./
RUN cargo build --release -p wifi-densepose-sensing-server \
    && strip target/release/sensing-server

FROM debian:bookworm-slim
COPY --from=builder /build/target/release/sensing-server /app/
COPY ui/ /app/ui/
EXPOSE 3000/tcp 3001/tcp 5005/udp
CMD ["/app/sensing-server --source auto --ui-path /app/ui --bind-addr 0.0.0.0"]

Built and pushed to GHCR:

docker build -f docker/Dockerfile.rust -t ghcr.io/zektopic/ruview:k8s-poc .
docker push ghcr.io/zektopic/ruview:k8s-poc
dns

The K8s Deployment

The deployment has one unusual requirement: UDP hostPort. The ESP32 sends raw CSI frames to a specific IP:port, so the pod needs to receive those packets directly on the host's network interface without kube-proxy NAT:

YAML
ports:
- containerPort: 5005
  protocol: UDP
  hostPort: 5005    # ESP32 sends directly to host

This means the pod must be pinned to a specific node (nodeName) — if it moves, the ESP32 would be sending to the wrong IP. For a homelab this is fine; for production you'd use a DaemonSet or a LoadBalancer with UDP support.

The HTTP API and WebSocket get standard NodePort services:

type: NodePort
ports:
- name: http
  port: 3000
  nodePort: 30900

An nginx reverse proxy ties it all together, handling WebSocket upgrades with proper timeout settings so the live data stream doesn't drop.

dashboard

What It Looks Like

The Observatory UI is the star — a cinematic Three.js dashboard with five holographic panels:

waves

Subcarrier Manifold

Live heatmap of all 56+ WiFi subcarriers, showing frequency effects.

favorite

Vital Signs Oracle

Breathing rate (6-30 BPM) & heart rate (40-120 BPM) from phase variations.

person_search

Presence Heatmap

Room-level signal field showing where people are located.

scatter_plot

Phase Constellation

Complex-plane plot of CSI phase, revealing movement patterns.

memory_alt

Convergence Engine

Signal processing pipeline metrics and overall health.

The Pose Fusion view goes further — it overlays WiFi-derived pose estimation onto a live camera feed. I connected my RTSP camera through Frigate's go2rtc, which already handles RTSP-to-HLS transcoding. The browser loads the HLS stream alongside the CSI data, and the fusion engine cross-correlates video motion with WiFi signal changes.

analytics

Real Data, Real Results

With the ESP32-S3 powered on and placed in my office, the system immediately detected:

  • sensors Presence: true with confidence ~0.78
  • directions_run Motion level: present_moving → present_still → active
  • group Person count: 1 (estimated from CSI subcarrier patterns)
  • speed 64 subcarriers streaming at 20 Hz

All through the wall, with no camera in the room.

JSON Response
{
  "classification": {
    "confidence": 0.78,
    "motion_level": "present_moving",
    "presence": true
  },
  "estimated_persons": 1,
  "features": {
    "breathing_band_power": 34.27,
    "motion_band_power": 61.17,
    "spectral_power": 158.92
  }
}
shield_lock
security

Privacy by Design

This is the compelling part. There is no camera in the sensing loop. The ESP32 captures WiFi signal disturbances — amplitude and phase changes caused by human bodies scattering radio waves.

There are no images, no video frames, no biometric data stored. The "sensing" is fundamentally different from surveillance.

For applications like elderly care monitoring, hospital patient tracking, or smart building occupancy — where cameras raise serious privacy and regulatory concerns — WiFi sensing sidesteps the problem entirely.

rocket_launch

What's Next

hub

Multi-node mesh

Adding 3-6 ESP32-S3 nodes for full 360-degree room coverage with multistatic fusion.

radar

ESP32-C6 + mmWave

Pairing the C6 with a Seeed MR60BHA2 60 GHz sensor for clinical-grade vital signs.

extension

Edge WASM modules

65 implemented edge intelligence modules run directly on the ESP32 as tiny WASM binaries (fall detection, sleep monitoring) with zero cloud dependency.

model_training

Training pipeline

Recording labeled CSI sessions to train the adaptive classifier for room-specific signal characteristics.

play_circle

Try It Yourself

The fastest path to a working system:

# 1. Docker (simulated data, no hardware)
docker run -p 3000:3000 ghcr.io/zektopic/ruview:k8s-poc
# Open http://localhost:3000/ui/

# 2. With ESP32-S3 hardware (~$9)
# Flash firmware, provision WiFi, run server with --source auto

# 3. Full K8s deployment
# See the deployment guide for complete instructions

The entire system — firmware, server, UI, signal processing, neural networks — is open source under MIT. One $9 microcontroller and some WiFi signals. That's all it takes to give a room spatial awareness.