Forum Replies Created
-
AuthorPosts
-
Hello Shengzhou,
As you found out, DPU is attached to the VM via pci-passthrough. You can perform installations of the DOCA Host and BF-Bundle packages, ultimately you can have access into the DPU over RShim.
Operating Mode: DPU mode – INTERNAL_CPU_OFFLOAD_ENGINE=ENABLED(0)
DPU↔NVMe PCIe topology: BF-3 functions and NVMe drives are under different PCIe roots and different CPU sockets
(Servers are configured with NPS=4)
- BF3 = PCI domain 0000:80, NUMA 7, socket 1
- NVMe = PCI domain 0000:20, NUMA 2, socket 0
I’m attaching 2 lspci outputs. Most of the DPUs are have the default PCI_SWITCH_EMULATION_ENABLE=0 whereas a few DPUs (HAWI, PSC, NCSA) are recently configured with PCI_SWITCH_EMULATION_ENABLE=1 for SNAP. You can see the PCI topology for both cases. Also, the topology does not provide the local PCIe peer path for direct HCA ↔ SSD P2P.
Firmware-config path: The only way to apply mlxconfig change (as well as activate a new firmware) seems to be cold-reboot of the host server (R7525). Level-3 (Driver restart and PCI reset) mechanism shows up as “supported” on the DPU, however it does not work to apply the NVCONFIG changes or the firmware activation from the DPU and ends up with “BF reset flow timeout”. Either the “known issues” on NVIDIA’s documentation for AMD servers, or the PCI-passthrough based setup cause issues. Cold-reboot of the host servers can be performed by the FABRIC Operations Team in coordination with the requests and active slivers on the servers.
CPU accounting: Slice vCPUs are not pinned as the default provisioning method. (CPU resources are not overcommitted – cpu_allocation_ratio=1). CPU pinning to the NUMA node of the selected device is supported by the control framework and can be utilized via node.pin_cpu(). I found the following links from API and an example.
- https://fabric-fablib.readthedocs.io/en/latest/node.html#fabrictestbed_extensions.fablib.node.Node.pin_cpu
- https://github.com/fabric-testbed/jupyter-examples/blob/main/fabric_examples/complex_recipes/iPerf3/iperf3_optimized.ipynb
BF-3’s 16 ARM cores are exclusive to the slice holding the card.
DPUs are available on 21 sites. I think the LoomAI tool has advanced filters that can show the availability as well as advance reservations. Or simply the following notebook can be used to query (with component types either “nic_connectx_7_100_available” or “nic_connectx_7_400_available”)
I’m not sure how these tools reveal the presence or availability, so I will note down here a complete set of sites that have DPUs.
- TACC, MICH, MASS, NCSA, WASH, DALL, SALT, UCSD, FIU, LOSA, NEWY, ATLA, SEAT, HAWI, KANS, RUTG, PSC, CERN, BRIST, AMST, TOKY
(One DPU per site, only WASH and SALT are 400G DPUs, others 100G)
We will look into the problem with bluefield.configure() if there is a fix needed.
On the other hand, dpu_ubuntu_24 image includes the BF-bundle file in /opt/bf-bundle directory. You can install according to standard installation steps (found in DOCA Downloads page)
Specifically for bluefield.config(), I believe it’s executing the following, so that you can issue manually.
sudo ip addr add 192.168.100.1/24 dev tmfifo_net0 sudo ip link set tmfifo_net0 up sudo bfb-install --bfb /opt/bf-bundle/<BFB-FILE> --rshim rshim0
Maintenance is completed. PSC node is online. VMs of the active slices are resumed.
September 9, 2026 at 7:14 pm in reply to: FABRIC HAWI – Management network connectivity issue #10039The problem is resolved. All VMs are accessible from the VM management network.
September 5, 2026 at 10:44 am in reply to: Active VM at SALT unreachable through management network #10034There was a power outage in the datacenter that caused all servers to reboot. Now, they are online. Their PCI devices are reattached (may require a reboot of the VM if they don’t show up).
Hello Ivan,
There is some corrections needed on the dataplane configuration which is affecting the SALT node for the DPU ports. Work is in progress and updates will be provided on this thread in the later hours today.
Best regards,
MertMaintenance completed.
August 16, 2026 at 1:25 am in reply to: Request for Host Cold Power Cycle to Apply BlueField-3 DOCA SNAP Firmware Config #9998Tanay,
There is a FABRIC ITSM ticket that you were included and you should have received emails from that. In case, you did not receive the emails from the ticket, I want also want to let you know over here, that DPUs on PSC, HAWI, NCSA are configured with
PCI_SWITCH_EMULATION_ENABLE=1. Please let us know if you have a chance to test for your application.OK, I found a VM (sliver) on TACC, that should be yours. That one shows a crashed kernel. 0fd2e355-ebaa-4372-9371-5a662a0e92cf-Node1-console
Without seeing the steps in your installation procedure, I don’t have an idea about this behavior. It can be helpful if you reveal your installation steps.
Hello Ivan,
Normally we use this forum “FABRIC Announcements” for announcements, but I think it’s fine to continue on this thread for a bit more, then we can all switch to another one under general questions.
If your slice is still active, can you send the Slice ID?
I will also share (here with everyone) the steps that I use for installing DOCA to the host and the DPU in the next hours.
Best regards,
MertI’m not sure about this but just in case I’m sharing. On the JupyterHub, from File > Hub Control Panel, “Stop My Server”, then “Logout” (top right corner), then login to JupyterHub again – may solve the problem with the token.
Hello Maureen,
I’m not sure what might be going on with the API part, but it can be better for us to understand if you share your view that shows the error.
Best regards,
MertHello Maureen,
Thank you for clarification. The fact that there are some filesystem actions with non-interactive execution on the current VM (some of which I indicated on my first comment) and previous VMs’ lifecycles along with the storage attachments are confusing to debug the root cause for the corrupted filesystem.
I took a snapshot of the storage volume, to try recovery attempts, but with the current look, I think it won’t be possible. You can reformat the volume any time.
Persistent Storage Volumes have been fine in general, but you can switch to other options, as Komal indicated CEPH-based distributed storage may be a better one. On the other hand, it should still be possible to replicate the data with multiple persistent storage volumes on different FABRIC sites, as you already have another one on WASH, and it’s possible to request more.
Sorry for the inconvenience, hopefully you can find other options on FABRIC that can better help with your work.
Best regards,
MertI’m looking at the previous access cycles to the volume. There seems to be multiple prior VMs (some of them with short lifespan) that attached/detached the volume on August 10th – roughly between 11am-4pm. I infer that the data in the volume could never have been accessed successfully recently (since last Friday) despite a VM remained attached during the weekend. Does that sound correct?
Hello Maureen,
We did not perform anything that could affect the persistent storage volumes during the maintenance. I wanted to take a look at the volume, but I’m seeing some events on journalctl output around 1:50am . Was there an attempt to reformat ?
Aug 13 01:49:03 fabric.rcnf sudo[112918]: rocky : PWD=/home/rocky ; USER=root ; COMMAND=/bin/file -s /dev/vdb Aug 13 01:49:03 fabric.rcnf sudo[112921]: rocky : PWD=/home/rocky ; USER=root ; COMMAND=/sbin/fdisk -l /dev/vdb Aug 13 01:49:18 fabric.rcnf sudo[113020]: rocky : PWD=/home/rocky ; USER=root ; COMMAND=/bin/cmp /dev/zero /dev/vdb -n 104857600 Aug 13 01:49:18 fabric.rcnf sudo[113023]: rocky : PWD=/home/rocky ; USER=root ; COMMAND=/bin/dd if=/dev/vdb bs=1M skip=500000 count=10 Aug 13 01:49:30 fabric.rcnf sudo[113056]: rocky : PWD=/home/rocky ; USER=root ; COMMAND=/sbin/mkfs.xfs /dev/vdb
-
AuthorPosts