1. Mert Cevik

Mert Cevik

Forum Replies Created

Viewing 15 posts - 1 through 15 (of 246 total)
  • Author
    Posts
  • in reply to: NVIDIA BlueField-3 DPUs #10056
    Mert Cevik
    Moderator

      Hello Shengzhou,

      As you found out, DPU is attached to the VM via pci-passthrough. You can perform installations of the DOCA Host and BF-Bundle packages, ultimately you can have access into the DPU over RShim.

      Operating Mode: DPU mode – INTERNAL_CPU_OFFLOAD_ENGINE=ENABLED(0)

      DPU↔NVMe PCIe topology: BF-3 functions and NVMe drives are under different PCIe roots and different CPU sockets

      (Servers are configured with NPS=4)

      • BF3 = PCI domain 0000:80, NUMA 7, socket 1
      • NVMe = PCI domain 0000:20, NUMA 2, socket 0

      I’m attaching 2 lspci outputs. Most of the DPUs are have the default PCI_SWITCH_EMULATION_ENABLE=0 whereas a few DPUs (HAWI, PSC, NCSA) are recently configured with PCI_SWITCH_EMULATION_ENABLE=1 for SNAP. You can see the PCI topology for both cases. Also, the topology does not provide the local PCIe peer path for direct HCA ↔ SSD P2P.

      Firmware-config path: The only way to apply mlxconfig change (as well as activate a new firmware) seems to be cold-reboot of the host server (R7525). Level-3 (Driver restart and PCI reset) mechanism shows up as “supported” on the DPU, however it does not work to apply the NVCONFIG changes or the firmware activation from the DPU and ends up with “BF reset flow timeout”. Either the “known issues” on NVIDIA’s documentation for AMD servers, or the PCI-passthrough based setup cause issues. Cold-reboot of the host servers can be performed by the FABRIC Operations Team in coordination with the requests and active slivers on the servers.

      CPU accounting: Slice vCPUs are not pinned as the default provisioning method. (CPU resources are not overcommitted – cpu_allocation_ratio=1). CPU pinning to the NUMA node of the selected device is supported by the control framework and can be utilized via node.pin_cpu(). I found the following links from API and an example.

      BF-3’s 16 ARM cores are exclusive to the slice holding the card.

      DPUs are available on 21 sites. I think the LoomAI tool has advanced filters that can show the availability as well as advance reservations. Or simply the following notebook can be used to query (with component types either “nic_connectx_7_100_available” or “nic_connectx_7_400_available”)

      I’m not sure how these tools reveal the presence or availability, so I will note down here a complete set of sites that have DPUs.

      • TACC, MICH, MASS, NCSA, WASH, DALL, SALT, UCSD, FIU, LOSA, NEWY, ATLA, SEAT, HAWI, KANS, RUTG, PSC, CERN, BRIST, AMST, TOKY
        (One DPU per site, only WASH and SALT are 400G DPUs, others 100G)
      in reply to: Bluefield Configuration Issue #10049
      Mert Cevik
      Moderator

        We will look into the problem with bluefield.configure() if there is a fix needed.

        On the other hand, dpu_ubuntu_24 image includes the BF-bundle file in /opt/bf-bundle directory. You can install according to standard installation steps (found in DOCA Downloads page)

        Specifically for bluefield.config(), I believe it’s executing the following, so that you can issue manually.

        sudo ip addr add 192.168.100.1/24 dev tmfifo_net0
        sudo ip link set tmfifo_net0 up
        sudo bfb-install --bfb /opt/bf-bundle/<BFB-FILE> --rshim rshim0
        
        
        in reply to: FABRIC PSC – Maintenance between Sept 8-14 #10045
        Mert Cevik
        Moderator

          Maintenance is completed. PSC node is online. VMs of the active slices are resumed.

          in reply to: FABRIC HAWI – Management network connectivity issue #10039
          Mert Cevik
          Moderator

            The problem is resolved. All VMs are accessible from the VM management network.

            in reply to: Active VM at SALT unreachable through management network #10034
            Mert Cevik
            Moderator

              There was a power outage in the datacenter that caused all servers to reboot. Now, they are online. Their PCI devices are reattached (may require a reboot of the VM if they don’t show up).

              in reply to: NVIDIA BlueField-3 DPUs #10009
              Mert Cevik
              Moderator

                Hello Ivan,

                There is some corrections needed on the dataplane configuration which is affecting the SALT node for the DPU ports. Work is in progress and updates will be provided on this thread in the later hours today.

                Best regards,
                Mert

                in reply to: FABRIC AMST – Maintenance on August 17 #10001
                Mert Cevik
                Moderator

                  Maintenance completed.

                  Mert Cevik
                  Moderator

                    Tanay,

                    There is a FABRIC ITSM ticket that you were included and you should have received emails from that. In case, you did not receive the emails from the ticket, I want also want to let you know over here, that DPUs on PSC, HAWI, NCSA are configured with PCI_SWITCH_EMULATION_ENABLE=1. Please let us know if you have a chance to test for your application.

                    in reply to: NVIDIA BlueField-3 DPUs #9997
                    Mert Cevik
                    Moderator

                      OK, I found a VM (sliver) on TACC, that should be yours. That one shows a crashed kernel. 0fd2e355-ebaa-4372-9371-5a662a0e92cf-Node1-console

                      Without seeing the steps in your installation procedure, I don’t have an idea about this behavior. It can be helpful if you reveal your installation steps.

                      in reply to: NVIDIA BlueField-3 DPUs #9995
                      Mert Cevik
                      Moderator

                        Hello Ivan,

                        Normally we use this forum “FABRIC Announcements” for announcements, but I think it’s fine to continue on this thread for a bit more, then we can all switch to another one under general questions.

                        If your slice is still active, can you send the Slice ID?

                        I will also share (here with everyone) the steps that I use for installing DOCA to the host and the DPU in the next hours.

                        Best regards,
                        Mert

                        in reply to: persistent storage attached but can’t mount #9989
                        Mert Cevik
                        Moderator

                          I’m not sure about this but just in case I’m sharing. On the JupyterHub, from File > Hub Control Panel, “Stop My Server”, then “Logout” (top right corner), then login to JupyterHub again – may solve the problem with the token.

                          in reply to: persistent storage attached but can’t mount #9983
                          Mert Cevik
                          Moderator

                            Hello Maureen,

                            I’m not sure what might be going on with the API part, but it can be better for us to understand if you share your view that shows the error.

                            Best regards,
                            Mert

                            in reply to: persistent storage attached but can’t mount #9981
                            Mert Cevik
                            Moderator

                              Hello Maureen,

                              Thank you for clarification. The fact that there are some filesystem actions with non-interactive execution on the current VM (some of which I indicated on my first comment) and previous VMs’ lifecycles along with the storage attachments are confusing to debug the root cause for the corrupted filesystem.

                              I took a snapshot of the storage volume, to try recovery attempts, but with the current look, I think it won’t be possible. You can reformat the volume any time.

                              Persistent Storage Volumes have been fine in general, but you can switch to other options, as Komal indicated CEPH-based distributed storage may be a better one. On the other hand, it should still be possible to replicate the data with multiple persistent storage volumes on different FABRIC sites, as you already have another one on WASH, and it’s possible to request more.

                              Sorry for the inconvenience, hopefully you can find other options on FABRIC that can better help with your work.

                              Best regards,
                              Mert

                              in reply to: persistent storage attached but can’t mount #9979
                              Mert Cevik
                              Moderator

                                I’m looking at the previous access cycles to the volume. There seems to be multiple prior VMs (some of them with short lifespan) that attached/detached the volume on August 10th – roughly between 11am-4pm. I infer that the data in the volume could never have been accessed successfully recently (since last Friday) despite a VM remained attached during the weekend. Does that sound correct?

                                in reply to: persistent storage attached but can’t mount #9978
                                Mert Cevik
                                Moderator

                                  Hello Maureen,

                                  We did not perform anything that could affect the persistent storage volumes during the maintenance. I wanted to take a look at the volume, but I’m seeing some events on journalctl output around 1:50am . Was there an attempt to reformat ?

                                   

                                  Aug 13 01:49:03 fabric.rcnf sudo[112918]: rocky : PWD=/home/rocky ; USER=root ; COMMAND=/bin/file -s /dev/vdb
                                  Aug 13 01:49:03 fabric.rcnf sudo[112921]: rocky : PWD=/home/rocky ; USER=root ; COMMAND=/sbin/fdisk -l /dev/vdb
                                  Aug 13 01:49:18 fabric.rcnf sudo[113020]: rocky : PWD=/home/rocky ; USER=root ; COMMAND=/bin/cmp /dev/zero /dev/vdb -n 104857600
                                  Aug 13 01:49:18 fabric.rcnf sudo[113023]: rocky : PWD=/home/rocky ; USER=root ; COMMAND=/bin/dd if=/dev/vdb bs=1M skip=500000 count=10
                                  Aug 13 01:49:30 fabric.rcnf sudo[113056]: rocky : PWD=/home/rocky ; USER=root ; COMMAND=/sbin/mkfs.xfs /dev/vdb
                                Viewing 15 posts - 1 through 15 (of 246 total)