Skip to content

fix: fall back to tracked residency when Vulkan iGPU free report wraps - #2124

Open
losewayy wants to merge 1 commit into
leejet:masterfrom
losewayy:igpu-capacity-accounting
Open

losewayy wants to merge 1 commit into
leejet:masterfrom
losewayy:igpu-capacity-accounting

Conversation

@losewayy

@losewayy losewayy commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

Summary

On integrated GPUs, treat an impossible Vulkan free-memory report (free > total, produced when the unsigned heapBudget - heapUsage subtraction wraps past the heap total) as unusable input rather than "no memory": fall back to the tracked-residency estimate total - resident that the same lambda already computes. Discrete GPUs keep the existing behavior — a wrapped report still maps to 0 available. Note that every non-Vulkan backend already treats this exact signature this way: a free > total report does not early-return for them — it falls through to the same min(free, total - resident) clamp. This change gives Vulkan iGPUs the same treatment, so it is parity with existing semantics rather than new policy. The fallback path also emits a one-line LOG_VERBOSE with the raw free/total values, so affected users can surface the wrapped report under -v for diagnosis.

Fixes #2022. May also unblock the same available 0.00 MB signature in #2073 if its failure was on the device-free side of check_capacity; if it was on the available_budget_bytes (--max-vram/auto-fit) side, this change does not apply — the two sides produce an identical log line.

Why

ggml_backend_vk_device_get_memory sums heapBudget - heapUsage per heap in unsigned arithmetic, and the Vulkan memory-budget spec allows usage to exceed the advisory budget — so a single over-budget heap wraps free to ~2^64. The Vulkan-only guard added in #2020 then reports available 0.00 MB and weight preparation aborts (cannot make enough memory available). On an iGPU this conclusion is particularly misleading: all heaps are backed by the same system RAM, the budget is a soft driver hint, and ggml's UMA allocator falls back to host-visible memory anyway — a wrapped figure is a broken report, not evidence of exhaustion.

When the report is impossible, the only trustworthy bound available here is the manager's own accounting, so the wrapped value is replaced by total and the existing min(free, total - resident) clamp applies — i.e., "what we believe is still free by our own books". This is a best-effort bound, not a guarantee of allocatable memory; genuine pressure still surfaces as an allocation failure instead of a wrong early abort. Nothing changes when the report is possible (wrapped or not, the discrete path is untouched).

Verification

Hardware: AMD Radeon 610M iGPU (uma: 1, two heaps ≈ 15.8 GiB shared pool) + NVIDIA RTX 5070 Ti Laptop (discrete Vulkan), Windows 11, master 2988060 + this change, SD_VULKAN=ON build.

  • The AMD Windows driver keeps heapBudget >= heapUsage, so the wrap is not naturally reachable here — measured with a standalone probe against ggml_backend_dev_memory: reported free decreases linearly across 0–15 GiB of device allocations (15.4 → 0.06 GiB) and allocations fail cleanly at the pool limit. Documented as a hardware/driver difference vs. the Intel report.
  • To exercise the branch deterministically I temporarily injected free = total + 1 KiB into available_device_bytes (local-only, reverted):
    • discrete Vulkan0: identical abort as reported — model manager cannot make enough memory available on Vulkan0: need 589.20 MB device / 77.20 MB budget, available 0.00 MB device / 3072.00 MB budget, failing at qwen_image_2_1 segment 1/34 (qwen_image_2_1.prelude);
    • iGPU Vulkan1 with the same forced report: all 34 weight-prep segments complete, sampling and VAE decode succeed, image saved (Qwen Image 2.1 Q6_K + Qwen3VL-8B Q4_K_M, --backend diffusion=vulkan1,vae=vulkan1,te=cpu --max-vram vulkan1=3 --diffusion-fa --vae-tiling, 256x256).
  • Regression runs (patched binary, no hook): iGPU reporter-mirror config and discrete-Vulkan generation both complete normally; free > total never occurs on these drivers, so the existing path is untouched there.

Not covered (intentionally out of scope)

  • Honest-but-misleading free reports on UMA (e.g., Linux GTT being allocatable while the device-local heap reads low, per @lvml's report in Vulkan on integrated GPU (UMA): model manager reports "available 0.00 MB device" — Qwen-Image 2.1 fails at weight prep #2022) need a different bound and are left for a follow-up. If desired, the iGPU fallback could be additionally clamped by available_ram_bytes() (currently in backend_fit.cpp) to account for untracked system-memory consumers — happy to add that as a follow-up or on request.
  • Drivers that classify a UMA/SoC part as eDiscreteGpu (e.g., some Strix Halo configurations) keep the old 0-return; covering them would need a different discriminator than the device type.
  • Other consumers of the raw free figure can also ingest the wrapped value — most notably ggml_graph_cut.cpp auto --max-vram detection, where a wrapped report effectively disables the auto limit, plus layer_split_partition.cpp capacities, diffusion_engine.cpp row split, and backend_fit.cpp planning. None of them abort on it, so this PR changes only the single aborting call site; normalizing all consumers behind a shared sanitizer is a reasonable follow-up.
  • The total - resident fallback is deliberately optimistic under genuine external pressure: on iGPUs total sums all heaps (which share the same physical RAM) and resident only counts manager-tracked bytes, so the estimate can exceed what is really allocatable when the OS or other apps hold memory. That failure stays on the normal allocation-error path — the Vulkan buffer-type allocator catches vk::SystemError and returns nullptr — rather than the early "cannot make enough memory available" abort. A host-RAM clamp (available_ram_bytes()) would tighten this; left as a follow-up to keep the diff minimal.
  • The per-heap saturating fix belongs in ggml (leejet/ggml) if preferred upstream; this change only makes sd.cpp robust to the impossible report.

Checklist

  • I have read and confirmed this PR follows the contribution guidelines.

Summary by CodeRabbit

  • Bug Fixes
    • Improved memory availability reporting for integrated GPUs using Vulkan when the reported free memory exceeds the total. This helps the app assess available device capacity more accurately in these cases.

@coderabbitai

coderabbitai Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 9dbfbb0f-d60d-4b6f-8a2a-c58b3776e87d


📥 Commits

Reviewing files that changed from the base of the PR and between 2988060 and cdc6677.



📒 Files selected for processing (1)
  • src/model_manager.cpp


Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 9 remain after this review.




📝 Walkthrough
📝 Walkthrough

Walkthrough

The Vulkan capacity check now handles reports where free memory exceeds total memory based on device type. Integrated GPUs use total memory as the free-memory value before tracked-residency accounting. Non-integrated GPUs continue to report zero available memory for this condition.

Changes

Vulkan memory accounting

Layer / File(s) Summary
Device-specific capacity check
src/model_manager.cpp
When a Vulkan report has more free bytes than total bytes, the check returns zero for non-integrated GPUs. For integrated GPUs, it logs the report and uses total bytes as free memory before the existing residency calculation.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~8 minutes

Change: Bug fix · Severity of issue fixed: Medium

Suggested reviewers: leejet



Merge Risk: ⚪ Minimal · up to cdc66

Integrated Vulkan GPUs with anomalous free-memory reports now use the existing residency estimate instead of reporting zero available memory; other devices retain their prior behavior. No actionable merge-blocking risk is established, so the change is ready for normal checks.

Architecture Summary

Architecture risk: 🔵 Low · up to cdc66

The change affects 1 system.

Changed systems: src

Architecture concerns
No architecture-level concerns identified.

Review details

Systems and components

  • observed — src (service) was modified; 1 changed file maps to changed impact.

Before / after behavior

  • observed — Modified behavior in src/model_manager.cpp: When Vulkan reports more free than total memory, the code now returns zero only for non-integrated GPUs. Integrated GPUs log the anomalous report and set free memory to the shared total, allowing the existing residency calculation to determine availability; previously every Vulkan device returned zero.


Pre-merge checks | Passed 4 | Failed 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 1 functions across 1 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check Passed The directly linked issue is #2022. The change handles the reported Vulkan UMA underflow signature: when free_bytes > total_bytes, an integrated GPU uses total_bytes with the existing tracked-resi…
Out of Scope Changes check Passed The reviewed change is limited to src/model_manager.cpp. It changes the memory-capacity decision for the #2022 underflow case and adds diagnostic logging for that case. No unrelated backend, allocat…
Title check Passed The title clearly and concisely describes the main change: using tracked residency when an integrated GPU reports wrapped Vulkan free memory.
Description check Passed The description provides a detailed summary, rationale, verification results, scope limitations, linked issues, and a completed checklist. The related issue and verification details appear under Summa…

  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR



  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Vulkan on integrated GPU (UMA): model manager reports "available 0.00 MB device" — Qwen-Image 2.1 fails at weight prep

1 participant