Spots

When GPU Resources Run Dry, It Still Looks Like Everything Is Working: A Taleโ€ฆ

๐Ÿ“ Originally published (in Japanese) at forge.workstyle.tech. ๐Ÿ“ Originally published (in Japanese) at forge.workstyle.tech. On the same machine (1 GPU, 12GB VRAM), two Claude Code sessions were running separate tasks concurrently. This one: Generating lip-sync caricature videos using InfiniteTalk (10โ€“12 minutes per video) The other: Estimating hand poses from real-life video footage (WiLoR) One day, the lip-sync task for 6 videos was still pending after an hour and a minute of waiting. Looks Normal on the Surface The GPU is running at full capacity. It looks like the slowness is due to heavy processing. I tried to break it down by process. In this environment (WSL2), per-process usage isn't visible. I have no idea who is using how much. It Was Happening on the Other Side Too In the other session, I asked about the speed of hand estimation.

Seven minutes per frame. And zero detections. The

Seven minutes per frame. And zero detections. The result was that when VRAM was running low, the hand estimation system was "successfully" returning the outcome that no hands were found. No errors were thrown. In 40 minutes, it only processed two frames, and both were empty.

Our lip-sync generation was also suffering, with VAE

Our lip-sync generation was also suffering, with VAE decode times dropping to 758 seconds per iteration. Under normal conditions, processing around 300 frames takes 10 to 12 minutes.

For both tasks, the GPU indicated it was

For both tasks, the GPU indicated it was running at "full load" the entire time it was exhausted, while virtually no progress was being made. Exhaustion Looks Deceptively Healthy

Normally, when things run slow, you check resource

Normally, when things run slow, you check resource utilization. If it's low, something is bottlenecked; if it's high, it's just chewing through a heavy workload.

Running out of VRAM is the exact opposite

Running out of VRAM is the exact opposite. Memory and utilization both peg at 100%, making the system look like it's firing on all cylinders. To make matters worse, one of them fails by returning a completely normal-looking "zero detections." You'd never catch it just by looking at the metrics. Comparison was made based on time required per unit, rather than display. Lip sync: 300 frames taking 10-12 minutes is normal. If it takes 1 hour, it's abnormal Hand estimation: 5ms per frame is normal. 431 seconds is 80,000 times slower By passing this data along with the start time to the other party, it's possible to determine "when it got stuck". We established the following rules for our sessions: Before starting a long GPU task, notify the other person (content, estimated VRAM usage, and duration). Wait while the other person is using the resources.

Once finished, let them know "I'm free." If

Once finished, let them know "I'm free." If you suspect a bottleneck, share the processing time per unit and the start time rather than just the display status.

From then on, the other person started notifying

From then on, the other person started notifying me in advance: "I'm going to use the GPU for a few minutes. VRAM will be 1-2GB, and it should take about 5-10 minutes." I also started reaching out before running batch processes of lip-sync generations. VRAM exhaustion can look "healthy" because both usage percentage and memory are maxed out. Depending on the inference process, failures might be returned normally as empty results. To distinguish issues, don't look at displaysโ€”look at the processing time per unit.

If you are sharing a GPU, notify others

If you are sharing a GPU, notify others before using it. If you're suspicious, pass the processing time and start timestamp. For further actions, you may consider blocking this person and/or reporting abuse

News

When GPU Resources Run Dry, It Still Looks Like Everything Is Working: A Tale of Two AI Competing for VRAM

๐Ÿ“ Originally published (in Japanese) at forge.workstyle.tech.

@spots #dev
Source: Dev.to
See more like this