System A display baseline — September 28, 2026¶
Historical baseline: every Seascape result in this report measures the existing Sway/Wayland presentation path. The standalone display probe presents generated test colors; its submission rates are not comparable application FPS. The subsequent component integration adds direct presentation to the real slicer and records its separate validation status. The matched direct-versus-Wayland results are reported separately. Do not interpret these baseline tables as direct-display measurements.
No production Suede code or installed binary was changed to collect these baseline measurements. Subsequent implementation is staged separately on System A; this historical report makes no improvement or rollout claim.
The retained System A baseline is the packaged NVIDIA 595.91.07 server driver. It is installed persistently and has already survived reboots; the saved 580 packages are rollback material, not an automatic reversion. Before the next hardware test, reboot to clear possible state from prior failing probes, verify driver/kernel/configuration and health, then warm up the workload. Capture the boot identity with each run. Compare Wayland and direct display on this same driver to separate backend benefits from the driver upgrade. Acceptance now explicitly requires reduced CPU/GPU overhead at equivalent frame rate, synchronization, and latency; reductions remain to be established with repeated full-pipeline measurements.
Identity and method¶
- Suede
v0.1.14-9-ge20a8f5; installed binary SHA-256:eae12d476f02563f70f2463fedd99a19fe357a25055c676edae59071ba1b9bfa. - Quadro RTX 8000, NVIDIA 580.178.04; Linux 7.0.0-34-generic; Sway 1.11.
- Four 1920×1080 outputs, reporting 59.939 Hz in Sway; direct scanout enabled.
- The running browser was Google Chrome 151.0.7922.71. The system API also lists installed Chromium 153; that is not the browser executable used by these baseline workloads. Executable hashes are in the raw captures.
- Current canvas: 900×536, aspect 1.68, existing warp and blend geometry. Separate loaded runs used render width 3200 with the same aspect/geometry.
- Each workload: 30-second warmup, one 60-second capture, approximately one-second CPU/GPU samples. Frame figures use only newly observed reports, excluding the initial potentially old stats snapshot.
- CPU is percent of one core from process tick deltas; 100% means one core. Means exclude unpaired process samples. These CPU windows are not claimed to align exactly with Suede's approximately ten-second stats reports.
Results¶
| Workload | Canvas width | Sway CPU | Slicer CPU | GPU utilization | Presented fps, mean | Straddles | Gate holds |
|---|---|---|---|---|---|---|---|
| As-found idle content | 900 | 0.17% | 0.56% | 0.0% | N/A | N/A | N/A |
| Sync, light source | 900 | 10.47% | 8.25% | 31.5% | 59.87 | 1 | 3 |
| Sync, Seascape source | 900 | 14.96% | 8.73% | 34.5% | 59.87 | 4 | 4 |
| Seascape content | 900 | 19.86% | 6.22% | 33.0% | 59.90 | 1 | 9 |
| Sync, Seascape source | 3200 | 19.27% | 10.30% | 55.3% | 59.55 | 17 | 20 |
| Seascape content | 3200 | 31.50% | 8.47% | 55.1% | 58.93 | 17 | 61 |
The idle run published no fresh frame intervals, consistent with damage-driven idle behavior. It is not a zero-FPS presentation failure. The three animated 900-pixel runs contain five, six, and five fresh intervals respectively; raw counts must be normalized by those intervals/frames before comparing arms. Their maximum absolute per-output phase was 0.0096, 0.0135, and 0.0116 ms. Every reported presentation on all four outputs used zero-copy scanout in those three runs. The larger canvas increased cost but did not saturate this GPU. These are loaded baselines, not measurements of a 99%-busy GPU.
Each condition has only one run; do not infer statistical confidence or a performance gain from these values. Repeat the matching conditions after any backend change and collect multiple alternating control/experimental runs. The original full configuration was restored and checked for exact equality after both the current-settings and larger-canvas matrices.
Wayland baseline on NVIDIA 595.91.07¶
The six-condition Wayland matrix was repeated on System A with NVIDIA package
595.91.07-0ubuntu0.26.04.2. The GPU remained a Quadro RTX 8000 and the kernel,
Suede binary, workload configuration, browser, and output setup matched the
580.178.04 captures. The running browser was Google Chrome 151.0.7922.71;
Chromium 153.0.8010.36 snap is separately listed by the system API and was not
the browser executable used for these workloads.
| Workload | Canvas width | Sway CPU | Slicer CPU | GPU utilization | Presented fps, mean | Straddles | Gate holds |
|---|---|---|---|---|---|---|---|
| As-found idle content | 900 | 0.17% | 0.34% | 0.03% | N/A | N/A | N/A |
| Sync, light source | 900 | 10.07% | 7.88% | 31.7% | 59.92 | 1 (5 windows) | 1 (5 windows) |
| Sync, Seascape source | 900 | 15.52% | 8.17% | 35.9% | 59.90 | 1 (6 windows) | 3 (6 windows) |
| Seascape content | 900 | 20.35% | 5.88% | 32.1% | 59.92 | 0 (5 windows) | 8 (5 windows) |
| Sync, Seascape source | 3200 | 14.81% | 8.53% | 55.1% | 59.87 | 1 (5 windows) | 3 (5 windows) |
| Seascape content | 3200 | 19.50% | 5.91% | 53.4% | 59.90 | 0 (5 windows) | 2 (5 windows) |
Each condition is one 60-second collection. CPU values are means of valid process-counter samples, reported as percent of one core; GPU utilization is the mean of available samples. FPS is the mean of fresh approximately ten-second stats windows. Straddle and gate-hold cells show raw counts followed by the number of fresh windows, so counts can be normalized across conditions. Idle again had no fresh frame intervals; its presentation metrics are unavailable, not zero. The configuration was restored and checked for exact equality after this matrix.
The active 3200-pixel sync-stress run averaged 207.95 W, 72.37 °C, and 1910.5 MHz, compared with 196.30 W, 67.23 °C, and 1919 MHz on 580.178.04. The 3200-pixel Seascape-stress run averaged 213.01 W, 76.28 °C, and 1905 MHz, compared with 208.35 W, 74.53 °C, and 1897.75 MHz. These are single-condition observations with differing thermal conditions, not evidence of a statistically established performance improvement. This section records the Wayland baseline only; the direct-display timing and fbdev experiments below are separate tests.
Raw JSON, the exact collector used, configuration backups, and probe logs are
kept locally in research/display-2026-09-28/ (ignored by Git because they
contain site information). The reproducible collection method and delegation
sequence are in the baseline runbook.
Research probes¶
An isolated headless-only Sway session, with a separate Chromium instance, provided 900×536 GPU DMA-BUF capture to the existing slicer at about 58.8 fps. This establishes capture feasibility, not daemon watchdog/readiness parity or production output planning under headless-only Sway.
Vulkan enumeration found four displays and four planes. Its mode list reports
59.940 Hz where Sway reports 59.939 Hz. The probe selects the explicit Vulkan
mode rather than assuming the two APIs use identical rounding. Vulkan display
names also differ from DRM/Sway connector names; connector IDs are mapped with
vkGetDrmDisplayEXT, not matched by display name.
On NVIDIA 580.178.04, System A exposed VK_KHR_present_wait/VK_KHR_present_id and
VK_EXT_display_control, but neither VK_GOOGLE_display_timing nor
VK_EXT_present_timing appeared in its capability dump. Completion waits and
vblank counters do not establish the per-image timestamp contract needed for
unchanged phase/lag statistics. Cross-display atomicity is also not guaranteed
by ordinary batched Vulkan presentation. See the primary references in
the corrected architecture assumptions.
The one-display test submitted 3,590 presents in 60.015 seconds (59.819/s)
and exited normally. The four-display test submitted 3,564 batches in 60.017
seconds (59.383/s per connector), then reported final-present completion on
all four connectors. It did not exit within the controller's 75-second
limit and was terminated by timeout (exit 124). The normal Sway and Suede
session restarted afterward. Instrumentation localized this to closing the DRM card after unloading the
Vulkan loader; every Vulkan destruction call had already returned. Keeping the
loader alive until after DRM close fixed the short four-display exit test. The
full-duration verification with that corrected ordering submitted 3,563 batches
in 60.009 seconds (59.375/s per connector), then completed every cleanup stage
and exited with status 0. This was a probe resource-lifetime bug, not an
established driver inability to release four displays.
Submission counts do not prove scanout FPS or phase.
An explicit SIGKILL of the four-display probe returned shell wait status 137. A new process immediately reacquired all four connectors and completed a five-second presentation run. Session recovery was controlled externally by the test harness; this does not implement or validate production slicer crash recovery under a direct backend.
Final restoration checks confirmed exact equality with the original saved
configuration, the original binary hash and cap_sys_nice=ep, four active
1920×1080 outputs at Sway's 59.939 Hz, the original app running, no divergences,
and all 17 health checks passing. No recovery timers remained pending.
The production backend remains unimplemented. The environment experiments below establish a usable timing source on 595, but not working present barriers. Neither fleet-wide support, synchronized cold start, nor production crash recovery is established by these standalone probes.
Environment follow-up¶
A read-only System B inventory found an RTX A1000 on NVIDIA 550.163.01. This
historical inventory compared against System A's then-installed 580.178.04 driver.
Both surveyed driver versions advertised VK_NV_present_barrier
but not VK_EXT_present_timing or VK_KHR_display_swapchain. No driver
was changed for this inventory.
NVIDIA's 595 release announcement
adds VK_EXT_present_timing. Its Vulkan driver release notes
also record a direct-display timing fix in developer driver 580.94.18.
Select a supported packaged driver containing this support; updating only the
Vulkan loader, tools, or SDK cannot add the driver's missing implementation.
Query actual surface timing stages and usable clock domains, then validate
per-image feedback before claiming phase/lag telemetry parity.
The present-barrier extension can synchronize corresponding presentation requests across swapchains; it is not limited to multiple GPUs. Query its device feature and each direct-display surface's support, then test it explicitly enabled. Advertisement alone is insufficient, and presentation synchronization does not establish atomic mode setting or physical scanout phase alignment.
System A is the primary test workhorse and System B is the production analog for confirmation. Test on System A first unless a documented hardware difference confounds the result. The recorded System B inventory above is historical; it does not imply tests or upgrades have been run there.
Completed System A experiments¶
The exact 580 server driver packages were saved for rollback before installing
Ubuntu's nvidia-driver-595-server=595.91.07-0ubuntu0.26.04.2. The kernel and
production Suede binary remained unchanged. The six Wayland workloads above
were repeated on 595 before testing direct display.
| Experiment | Observed result |
|---|---|
| Present barrier, 580.178.04 | Extension, device feature, and all four surfaces report support; the first batched present fails with ERROR_UNKNOWN for connector 141. |
| Present barrier, 595.91.07 | Same advertised support and same failure, with and without the validation layer. |
| Present timing, 595.91.07 | VK_EXT_present_timing revision 3 and VK_KHR_present_id2 are supported. All four surfaces return complete per-present timing records. |
| Shared display swapchains, 595.91.07 | VK_KHR_display_swapchain remains absent. A shared input canvas remains possible; shared scanout allocation is a separate capability. |
Temporary nvidia_drm.fbdev=0, 595.91.07 |
The short timing and barrier tests produced no kernel flip warnings/timeouts, unlike the original fbdev setting. The barrier still failed. Original boot settings were restored afterward. |
The timing stage used is IMAGE_FIRST_PIXEL_OUT, not
IMAGE_FIRST_PIXEL_VISIBLE. System A exposes DEVICE and local presentation clock
domains, but no direct CLOCK_MONOTONIC presentation domain. Its
timestampPeriod is 1 ns, allowing DEVICE presentation nanoseconds to be mapped
to host monotonic time using VK_KHR_calibrated_timestamps. The probe records
calibration pairs about every 250 ms and rejects unbracketed events or pairs
whose clock-rate mismatch exceeds its consistency threshold. This is an
estimate with separately reported calibration deviation and observed mismatch,
not a proven error bound between samples.
A single start/end calibration over ten seconds drifted by 41,902 ns and was correctly rejected. Periodic calibration in the fbdev-off ten-second test mapped all 2,224 returned records (556 per output); all 38 adjacent sample pairs passed the consistency checks. Maximum reported calibration deviation was 2,944 ns and maximum adjacent-pair mismatch was 456 ns. There were no missing, duplicate, zero, or incomplete records. The largest equal-present-ID span during startup was approximately 571 ms. After excluding the first two seconds relative to the latest output's first timestamp, the maximum span was 1,280 ns. Those measurements establish useful software feedback, not synchronized cold start or physical genlock. The clear-color probe does not exercise the full capture, warp, blend, and settle pipeline.
Final verification on the retained configuration¶
With the original fbdev=1 setting restored and validation disabled, the
60-second timing run submitted 3,562 batches in 60.007 seconds. Each of the
four outputs returned exactly IDs 1–3,562: 14,248 complete, nonzero records,
with no missing or duplicate IDs. All 238 adjacent calibration pairs passed
the consistency checks, mapping every record. Maximum reported calibration
deviation was 11,328 ns; maximum adjacent interval mismatch was 2,460 ns.
The whole-run clock mismatch was 195,419 ns, reinforcing the need for periodic
calibration rather than a single fixed offset.
The maximum equal-ID span was 457.6 ms during startup and 1,216 ns after the
same two-second exclusion used above. The software evidence therefore does
not satisfy the synchronized-cold-start gate. Normal teardown returned
success but took 13.112 seconds; total probe wall time was 77.27 seconds.
Kernel flip warnings and per-head flip timeouts returned with fbdev=1.
Automatic recovery restored the normal service. Validation-enabled tests had
reported no Vulkan validation errors; this does not rule out application or
driver defects.
System A retains NVIDIA 595.91.07 and the original kernel. Final checks confirmed
exact restoration of the original Suede configuration and boot configuration
files, unchanged installed binary hash and capabilities, four active
1920×1080 outputs at 59.939 Hz, the original application running, no
divergences, all 17 health checks passing, and no pending display-recovery
timers. Exact 580 driver rollback packages remain saved on System A at
/var/tmp/suede-display-driver-20260928/rollback-packages/. System B was not
modified or used for these experiments.
The next gate is synchronized startup and reliable recovery on System A, followed by integration of this timing contract with Suede's existing settle logic. The repeated barrier failure is now explained; see the root cause section. An alternate synchronization design is needed before claiming atomic wall updates. Full headless daemon readiness/watchdog behavior also remains unverified. Confirm a viable path on System B afterward; use frame-coded camera observations to check wall coherence independently of software feedback. CPU cost, GPU work, power, and throughput remain measurable regardless of timestamp availability.
System B projector comparison¶
The user's paired-projector inventory justified a hardware comparison after the System A experiments. System B remained on its existing RTX A1000, NVIDIA 550.163.01, and Linux 6.12.101+deb13-amd64. No driver, boot setting, production binary, or saved Suede configuration was changed.
Only three projectors were connected during the test: DP-5 (DRM connector 89), DP-7 (97), and DP-8 (101). Configured output DP-6 (93) was already disconnected. All active projectors used 1920×1200 at 59.950 Hz. EDID identifies DP-7 and DP-8 as the same product, with identical detailed 1920×1200 timing descriptors. DP-5 advertises the same 154 MHz pixel clock and 2080×1235 totals, but different sync-polarity flags. The missing projector prevented a four-output or second-identical-pair test.
The same standalone probe was run without validation, with five-second presentation intervals and independent timed session recovery:
| Test | Result |
|---|---|
| Ordinary presentation, all three | 292 batches in 5.003 s; final-present waits completed on all heads; normal exit 0. |
| Barrier, all three | First present failed with ERROR_UNKNOWN, reported for connector 101; exit 3. |
| Barrier, identical DP-7/DP-8 pair | First present failed with ERROR_UNKNOWN, reported for connector 101; exit 3. |
| Barrier, DP-7 alone | First present failed with ERROR_UNKNOWN, reported for connector 97; exit 3. |
Every requested surface, along with the device feature and extension, reported
barrier support, and barrier-enabled swapchain creation succeeded. The
ordinary control's normal teardown took 609 ms, with total wall time 7.816 s.
These short runs do not establish sustained performance or synchronization;
550 does not expose VK_EXT_present_timing. Different GPU, driver, kernel,
output count, and display mode prevent attributing the teardown difference
from System A to any single factor.
Kernel logs also contain nv_drm_revoke_modeset_permission warnings, including
during the initial inventory before the presentation matrix. They are not
specific evidence of a barrier failure; recovery checks below concern the
restored application and output state, not absence of kernel warnings.
Failure with an identical pair makes mixed projector models a weaker explanation; the single-projector failure also warrants a minimal feature interaction/control test. The barrier runs followed each other in the same session outage, so retained driver state after a prior failure is not excluded. This evidence does not establish whether the defect is in the application, driver, or required system setup.
Automatic recovery restored all three projectors and the original application. Checks confirmed exact saved-configuration equality, unchanged binary hash and capabilities, unchanged driver, original active output modes, no new configuration divergences, no failed health checks, and no pending recovery timer. The pre-existing disconnected DP-6 and video-decode warnings remain; the initial phase warning was 1.1 ms and can change when the session restarts. Raw logs, EDIDs, and before/after snapshots are in local, ignored research notes for System B. Controller setup attempts that stopped before running a probe are preserved separately from test logs.
Relationship to the current Sway path¶
Suede already coordinates outputs above Wayland: in locked mode, it waits for
every participating output's prior wp_presentation outcome before submitting
the next wall update, and uses the same completed capture across the slices.
An unresponsive output is eventually excluded from the gate so it cannot
freeze the wall. Presentation feedback reports a realized presentation event;
it is distinct from a frame callback inviting the client to draw again. See
the protocol source
and the current synchronization description.
Separately, Suede batches Sway output setup and can disable/re-enable all heads together to improve startup phase alignment. The measured improvement on NVIDIA is not a portable physical clock-lock guarantee. The feedback gate prevents outputs from continually running ahead; it cannot retroactively prevent a frame mismatch caused by one head missing its presentation deadline. A present barrier coordinates the new presentation requests inside the driver.
Therefore working VK_NV_present_barrier is a candidate for stronger
coordination, not a feature the current Wayland path already depends on.
A direct backend could seek parity by preserving Suede's existing gate with
validated completion/timing feedback and independently solving output setup
and recovery. It must demonstrate that parity under load and faults before
replacing the current path; a failed vendor barrier alone does not prove that
parity is impossible.
Seascape resolution sweep¶
Presentation backend: existing Sway/Wayland, control arm only. The direct display arm still requires integration with Chrome capture and Suede's existing warp/blend renderer before this sweep can be repeated under that path.
System A was rebooted before this sweep, which ran September 29, 2026 UTC
(September 28 in the user's time zone). The new boot ID was
39da87be-9c55-4e30-b90d-1aa55487e656. NVIDIA 595.91.07, kernel
7.0.0-34-generic, the installed Suede binary identified above, and Google Chrome
151.0.7922.71 were retained. Health passed 17/17 after boot.
The existing Seascape application ran without a Suede test-pattern overlay.
Only canvas renderWidth varied between conditions; the 1.68 aspect, four
1920×1080 physical outputs at 59.939 Hz, warp/blend geometry, and remaining
configuration stayed fixed. Every point used a temporary preview with
revision/generation/epoch preconditions and independent timed rollback, followed
by exact restoration of committed revision 1979. No production binary changed.
The live Seascape page sets its WebGL canvas dimensions from the window size
times device pixel ratio and draws a fullscreen shader each animation callback.
Increasing canvas size therefore increases rendering work, rather than merely
scaling a fixed-size shader image. A source snapshot is saved with SHA-256
73fb187dc11ee3ac2728cababb74a0b593b9fc20827ece0b79c71d7b9023e114.
Each condition had a 30-second warmup and a 60-second collector run. The order
was 3200, 4500, 4000, 6000, 4500, 3200, 4500 pixels wide.
| Canvas | Run | Presented fps, mean | GPU utilization, mean | Total CPU, % of one core | GPU temperature, °C | Graphics clock, MHz |
|---|---|---|---|---|---|---|
| 3200×1905 | 1 | 59.89 | 54.47% | 73.74% | 67.45 | 1917 |
| 4000×2381 | 1 | 59.84 | 70.65% | 56.45% | 79.13 | 1868 |
| 4500×2679 | 1 | 51.23 | 74.82% | 75.92% | 77.73 | 1849 |
| 6000×3571 | 1 | 35.88 | 91.62% | 75.19% | 81.77 | 1772 |
| 4500×2679 | 2 | 57.42 | 82.08% | 71.87% | 81.10 | 1816 |
| 3200×1905 | 2 | 59.28 | 56.22% | 67.90% | 76.98 | 1889 |
| 4500×2679 | 3 | 50.21 | 76.32% | 74.99% | 80.48 | 1837 |
FPS is the mean of fresh approximately ten-second presentation intervals; 4500 run 1 contains six intervals and every other run contains five. The first potentially stale snapshot is excluded. Total CPU sums Sway, slicer, daemon, and browser for each sample only when all four roles have valid tick deltas; partial totals are excluded. GPU figures are utilization samples, not measured shader execution time. All seven captures completed without collector errors, and their executable hashes, driver, kernel, physical output configuration, bootstrap, and compositor environment matched. Configuration differed only in canvas width across the measured Seascape conditions.
Use 4500×2679 as the primary workload for testing recovery of frame rate. All three runs failed to sustain the 59.939 Hz output rate, with means spanning 50.21–57.42 fps. Their lowest fresh-window rates were 49.63, 55.51, and 48.39 fps. The three runs recorded 33, 36, and 44 straddles, and 353, 153, and 283 gate holds respectively; these are raw counts over six, five, and five intervals and must be normalized before comparing arms. Retain 4000×2381 as a near-full-rate case and 6000×3571 as a heavier overload case; the latter has only one exploratory run. The warmer 3200 control recovered close to full rate but had some frame loss, so it is not an invariant zero-drop reference.
The variable 4500 results are real run variation, not an established backend effect. Seascape animates its camera/scene, GPU temperature and clock changed, and the driver/pipeline can vary; these data do not identify the cause of the difference. Repeat and alternate control/experimental runs on the same driver before claiming improvement. Higher achieved FPS at the same workload is a throughput result; lower overhead at equivalent FPS is a separate comparison.
The captured canvasFps is the slicer's capture rate, not a direct measurement
of Chrome's rendered FPS. Likewise, perFrameMs.gpu measures host wall time
waiting for the Vulkan fence, not GPU timestamp execution time. This sweep
establishes a demanding Chrome workload that overloads the existing complete
pipeline; it does not isolate Chrome's shader as the sole bottleneck.
Final verification confirmed the original 900×536 canvas and active application,
all four physical output modes, exact committed configuration, unchanged Suede
hash/capabilities and driver, no divergences, 17/17 passing health checks, and
no pending resolution-recovery timers. Raw captures, summaries, controller
logs, and the complete remote evidence archive are under the ignored
research/display-2026-09-28/resolution-sweep/ directory; the source machine
also retains /var/tmp/suede-resolution-sweep-20260929/.
Present barrier root cause — September 29, 2026¶
The barrier failure was traced with an LD_PRELOAD shim that logs every
nvidia-drm, NVKMS (/dev/nvidia-modeset), and Resource Manager ioctl the
NVIDIA user-space driver issues from the probe process, cross-checked against
the NVKMS sources in open-gpu-kernel-modules at tag 595.91.07. System A stayed
on 595.91.07 and kernel 7.0.0-34 throughout; every run below used the standard
harness (Suede and the tty1 session stopped, automatic recovery afterwards).
- The driver implements
VK_NV_present_barrierwith an NVKMS swap group. On the first barrier present it issues fourSET_MODEcalls and thenNVKMS_IOCTL_ALLOC_SWAP_GROUP, which returnsEPERM; the driver reports that asVK_ERROR_UNKNOWNon the last swapchain of the batch. NVKMS only allows swap groups, flip-lock groups, and framelock attributes for the modeset owner or sub-owner (nvKmsOpenDevHasSubOwnerPermissionOrBetter). A client that acquired its displays throughvkAcquireDrmDisplayEXTnever is one: nvidia-drm grants the driver's NVKMS handle per-head modeset permission only. This is independent of display count, model, and driver version, which matches the System B single-projector and identical-pair failures. - Sub-ownership can be forced from inside the process. The DRM master may
issue
DRM_IOCTL_NVIDIA_GRANT_PERMISSIONS(SUB_OWNER)with a fresh NVKMS token, and the process can then callNVKMS_IOCTL_ACQUIRE_PERMISSIONSon the driver's own/dev/nvidia-modesethandle, found through/proc/self/fd. The grant must follow display acquisition (with full permissions already held,vkAcquireDrmDisplayEXTfails withVK_ERROR_INITIALIZATION_FAILED) and must be revoked before display release, because nvidia-drm rejects every atomic commit while it stands. With the grant,ALLOC_SWAP_GROUPandSET_SWAP_GROUP_CLIP_LISTsucceed. - The next step needs Quadro Sync hardware. The driver then sets the disp
attribute
FRAMELOCK_SYNC = 0, which NVKMS rejects when the GPU has no framelock device (SetFrameLockSyncreturns false whenpFrameLockEvois null). When the shim reports success for that call, the driver registers a surface and a deferred request FIFO, queries the first display's dynamic data, and fails in user space with no further kernel call; the failure strings inlibnvidia-eglcoresit besideQuadroSyncServerDpy,QuadroSyncClientDpys, andQuadroSyncHouseOutputregistry-key parsing. NVKMS's own flip-lock group ioctl is a no-op for one GPU (EnableLockGroupFlipLock: "TODO: enable fliplock for single GPUs"), so no kernel path offers single-GPU cross-head flip lock outside swap groups. - No driver release changes this. The same permission check, the same framelock early return, and the same single-GPU TODO are present at tag 615.71.09, the newest published kernel-module source (September 2026). The 595.44.15 Vulkan beta and the 610 and 615 branches list no barrier or direct-display changes, and 615.71.09 carries a separate regression of three-second blocking commits per head on compositor exit. System A's Quadro RTX 8000 accepts a Quadro Sync II board; System B's RTX A1000 has no sync connector, so a hardware barrier is not available on the production analog.
Sub-ownership as a fix for teardown and kernel warnings¶
Sub-ownership also switches nvidia-drm into a passive mode: it stops handling
NVKMS flip events (the source of the nv_flip == NULL warnings) and does not
issue its own blocking commits while the grant stands. A matched pair of
ten-second --present-timing probe runs on the retained fbdev=1 boot gave:
| Run | Kernel warnings | Flip event timeouts | Teardown | Probe wall | Process CPU (user / sys) |
|---|---|---|---|---|---|
| Per-head modeset permission | 4 | 4 (3 s each) | 13,183 ms | 27.2 s | 0.35 s / 6.81 s |
| NVKMS sub-ownership | 0 | 0 | 929 ms | 15.2 s | 0.32 s / 7.02 s |
The presenter now grants sub-ownership after acquiring its displays, revokes
it before releasing them (also on the unverified-work path), and clears a
stale grant left by a crashed predecessor at startup; the probe exposes the
same behavior as --nvkms-sub-owner and --revoke-sub-owner, and the System A
harness's recovery script revokes before restarting the session. The grant is
issued only when DRM_IOCTL_VERSION names nvidia-drm. Because nvidia-drm
refuses atomic commits while the grant stands, a compositor cannot light the
outputs until it is revoked; a crashed direct renderer therefore needs the
startup cleanup or the probe's revoke mode before Sway can present again.
Where the direct path's CPU goes¶
Per-thread /proc sampling of the probe shows two different costs. About
4.8 s of kernel CPU accrues in the first five seconds, before any present:
the driver's display enumeration issues 163 QUERY_DPY_DYNAMIC_DATA calls of
up to 95 ms each, and perf attributes the busy time to one Resource Manager
function entered through ioctl. Steady-state presenting of four
1920×1080 heads then costs about 17% of one core in kernel time plus 2.5% in
user time. Per-ioctl timing during 7.4 s of presenting (442 batches) shows:
| Kernel call per present batch | Count | Total | Mean |
|---|---|---|---|
RM object allocation, class 0x0005 (NV01_EVENT) |
442 | 320 ms | 0.72 ms |
RM object free (NV_ESC_RM_FREE) |
452 | 336 ms | 0.74 ms |
NVKMS_IOCTL_FLIP (one per head) |
1,769 | 173 ms | 0.10 ms |
The allocate/free pair is issued once per vkQueuePresentKHR, not per
swapchain, so batching already amortizes it. It is unchanged without present
timing, without present IDs, and with three swapchain images instead of two,
so it is intrinsic to the driver's direct presentation and no application-side
change removes it.
The cost is a GSP artifact. With GSP firmware enabled (the 595 default on
Turing), each Resource Manager object allocation and free is a synchronous
round trip to the GPU's system processor. System A was rebooted with
NVreg_EnableGpuFirmware=0 (/etc/modprobe.d/nvidia-gsp-off.conf, retained;
the proprietary kernel module supports this on Turing, the open module does
not) and the same probe runs were repeated:
| Measure, four heads at 59.94 Hz | GSP on | GSP off |
|---|---|---|
| Kernel time in ioctls per 7.4 s of presenting | 845 ms | 267 ms |
| RM allocation per batch, mean | 0.72 ms | 0.05 ms |
| RM free per batch, mean | 0.74 ms | 0.12 ms |
NVKMS_IOCTL_FLIP per head, mean |
0.10 ms | 0.10 ms |
| Steady-state kernel CPU, one core | about 17% | about 4.6% |
| Steady-state user CPU, one core | about 2.5% | about 2.7% |
| Submitted batches per second | 56.9 | 57.0 |
| Probe teardown (sub-ownership) | 929 ms | 790 ms |
The one-off enumeration cost at startup is unchanged. These are standalone probe measurements on one boot each; the full pipeline, and the Wayland path on the same boot, still need the repeated matched runs the measurement contract requires before any appliance-level claim. The boot also produced no nvidia-drm flip warnings or timeouts across the sub-ownership runs.
One real-slicer direct run through the integration harness (four heads, 4500×2679 Seascape, 30 s, GSP off) confirmed the presenter change end to end: the slicer logged the grant and the revoke, exited normally 0.7 s after its run time (the earlier r2 runs needed about 13 s of teardown), the kernel log stayed clean, and the appliance restored with 17/17 checks. Its second stats interval reported 53.2 fps presented with 6 straddles; that is a single unwarmed run and not a performance result.