# MonaiLabel CUDA out of memory with Bundle

**URL:** <https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914>\
**Category:** Support\
**Tags:** segmentation, monailabel\
**Created:** [June 8, 2023, 10:09am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914 "2023-06-08T10:09:01Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![FloCo](https://avatars.discourse-cdn.com/v4/letter/f/2bfe46/32.png) [@FloCo](https://discourse.slicer.org/u/FloCo)\
**Post date:** [June 8, 2023, 10:09am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/1 "2023-06-08T10:09:01Z")

</div>

Hello everyone,

I am working with Slicer to do machine learning using MonaiLabel, the annotation and learning part work perfectly, it’s really nice tool !  
Unfortunately I have troubles when I run the automatic segmentation with medium and big CT scan (starting from around 256_256_256 pixels I would say).  
I got an error message :  
“torch.cuda. OutfMemoryError: CUDA out of memory. Tried to allocate 7.28 GiB (GPU 0; 23.65 GiB total capacity; 12.28 GiB already allocated; 2.67 GiB free; 19.48 GiB reserved in total by PyTorch”

I have a NVIDIA GeForce RTX 4090: 24GB memory, and I am using MonaiLabel -Zoo/Bundle (0.7.0rc6), note that I have tried other versions of Bundles with same result but if I use Radiology application it works well.

Since the training is working perfectly and since I have 24GB, I would have thought that it would be ok for CT scan of medium sizes.

Do you have any ideas if I do something wrong about the memory ?

Thank you for your help  
Florian COTTE

---

<div class="post-metadata">

**Author:** ![ylcnkzy](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.slicer.org/ylcnkzy/32/12189_2.png) [@ylcnkzy](https://discourse.slicer.org/u/ylcnkzy)\
**Post date:** [June 9, 2023, 6:47am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/2 "2023-06-09T06:47:36Z")

</div>

Hi Florian,  
Have you checked if you have the recent version of Cuda (or at least the cuda version you have supports RTX 4090)?

It is better to check it [here](https://docs.nvidia.com/deploy/cuda-compatibility/)

Best

---

<div class="post-metadata">

**Author:** ![FloCo](https://avatars.discourse-cdn.com/v4/letter/f/2bfe46/32.png) [@FloCo](https://discourse.slicer.org/u/FloCo)\
**Post date:** [June 9, 2023, 10:22am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/3 "2023-06-09T10:22:00Z")

</div>

Hi ylcnkzy,

I have cuda version 11.8 and from what I have found on internet it seems that it supports well RTX4090.  
I can try to see if it’s a problem with the driver, I’ll look into that.

Best,

---

<div class="post-metadata">

**Author:** ![rbumm](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.slicer.org/rbumm/32/9404_2.png) [@rbumm](https://discourse.slicer.org/u/rbumm)\
**Post date:** [June 9, 2023, 11:11am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/4 "2023-06-09T11:11:41Z")

</div>

Can @diazandr3s weigh in? Would also be curious about the boundaries.

---

<div class="post-metadata">

**Author:** ![diazandr3s](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.slicer.org/diazandr3s/32/9973_2.png) [@diazandr3s](https://discourse.slicer.org/u/diazandr3s)\
**Post date:** [June 10, 2023, 3:17am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/5 "2023-06-10T03:17:20Z")

</div>

Thanks for the ping @rbumm.

@FloCo which model are you using to run inference?

Here is the GPU memory usage for the two available models (1.5mm and 3mm): [model-zoo/models/wholeBody\_ct\_segmentation at dev · Project-MONAI/model-zoo · GitHub](https://github.com/Project-MONAI/model-zoo/tree/dev/models/wholeBody_ct_segmentation#gpu-consumption-warning)

---

<div class="post-metadata">

**Author:** ![FloCo](https://avatars.discourse-cdn.com/v4/letter/f/2bfe46/32.png) [@FloCo](https://discourse.slicer.org/u/FloCo)\
**Post date:** [June 10, 2023, 5:10am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/6 "2023-06-10T05:10:14Z")

</div>

Thank you for your response @diazandr3s , indeed I didn’t see this page on memory usage.  
I looked into totalsegmentator ressources and it was much less, it’s not the same model ?

> **[GitHub - wasserth/TotalSegmentator: Tool for robust segmentation of 104...](https://github.com/wasserth/TotalSegmentator)**
>
> Tool for robust segmentation of 104 important anatomical structures in CT images - GitHub - wasserth/TotalSegmentator: Tool for robust segmentation of 104 important anatomical structures in CT images

I am using the model with 1.5mm so everything is normal then.

I have 2 GPUs 4090 of 24GB, is there a way a using both for inference and doing like one of 48GB ?

Thank you

---

<div class="post-metadata">

**Author:** ![rbumm](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.slicer.org/rbumm/32/9404_2.png) [@rbumm](https://discourse.slicer.org/u/rbumm)\
**Post date:** [June 10, 2023, 5:49am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/7 "2023-06-10T05:49:48Z")

</div>

No, it is not the same model, but wholeBody\_ct\_segmentation was trained using TotalSegmentator data.

---

<div class="post-metadata">

**Author:** ![FloCo](https://avatars.discourse-cdn.com/v4/letter/f/2bfe46/32.png) [@FloCo](https://discourse.slicer.org/u/FloCo)\
**Post date:** [June 10, 2023, 6:16am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/8 "2023-06-10T06:16:57Z")

</div>

Oh I see ! My mistake then .  
And about the fact that I have two GPUs ? I guess I can’t “merge” them but I want to be sure I am not missing something simple to do ?

Thank you

---

<div class="post-metadata">

**Author:** ![rbumm](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.slicer.org/rbumm/32/9404_2.png) [@rbumm](https://discourse.slicer.org/u/rbumm)\
**Post date:** [June 10, 2023, 9:09am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/9 "2023-06-10T09:09:02Z")

</div>

It seems to be possible to use two GPUs, although I never tried this, [see here](https://github.com/Project-MONAI/MONAILabel/discussions/1400). The question is whether it has any impact on the CUDA Out of memory error. You could probably add a MONAILabel issue on GitHub or comment in the thread above.

---

<div class="post-metadata">

**Author:** ![curtislisle](https://avatars.discourse-cdn.com/v4/letter/c/48db29/32.png) [@curtislisle](https://discourse.slicer.org/u/curtislisle)\
**Post date:** [June 10, 2023, 2:11pm UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/10 "2023-06-10T14:11:12Z")

</div>

There is complexity here, and I am not expert, but in my experience, pytorch DNN models can be configured to use two GPUs instead of one. I assume many MONAI models take advantage of this. Use of two GPUs generally divides memory requirements in half. In the link Rudolf posted, it shows how to set the available devices to include 0 and 1. This is necessary. It might be the configuration of MONAILabel does this for us. I could test on a two GPU server in two weeks if we don’t have an answer earlier.

---

<div class="post-metadata">

**Author:** ![curtislisle](https://avatars.discourse-cdn.com/v4/letter/c/48db29/32.png) [@curtislisle](https://discourse.slicer.org/u/curtislisle)\
**Post date:** [June 10, 2023, 4:44pm UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/11 "2023-06-10T16:44:17Z")

</div>

After thinking about this. Using two GPUs halves memory only if the batch size for training or inferencing is greater than 1. From the MONAILabel GitHub page, it looks like the 1.5mm totalsegmentator model requires more than 24G memory for some volumes. You  
Might split your volume into two sets of slices and separately segment each half then combine the segmentations. You would want to examine the boundary to look for discontinuities.

Combining the memory of two GPUs into one has been done by NVIDIA but requires custom work. Hope this helps.

---

<div class="post-metadata">

**Author:** ![FloCo](https://avatars.discourse-cdn.com/v4/letter/f/2bfe46/32.png) [@FloCo](https://discourse.slicer.org/u/FloCo)\
**Post date:** [June 11, 2023, 2:25pm UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/12 "2023-06-11T14:25:38Z")

</div>

I was thinking about splitting dicom images in two but I am not very confident about the continuity, I can try this and see how it goes.

About merging merging memory of GPUs, it sounds that it would solve the problem on my case, I am not an expert so it might be too difficult for me but I’ll follow this idea too thank you !

---

<div class="post-metadata">

**Author:** ![xiang\_yao](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.slicer.org/xiang_yao/32/68673_2.png) [@xiang\_yao](https://discourse.slicer.org/u/xiang_yao)\
**Post date:** [August 1, 2024, 2:29am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/13 "2024-08-01T02:29:29Z")

</div>

Can multi GPU inference be used on Windows or Linux systems? Currently, due to data size issues, it is evident that monailabel infers data in memory (i.e. CPU) without utilizing CUDA acceleration.

If I have multiple 4090 cards on a computer, how can I involve them in parallel computing for inference?

Have you solved the problem of merging the memory of two graphics cards mentioned above?

---

<div class="post-metadata">

**Author:** ![ece](https://avatars.discourse-cdn.com/v4/letter/e/3be4f8/32.png) [@ece](https://discourse.slicer.org/u/ece)\
**Post date:** [August 1, 2024, 2:50am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/14 "2024-08-01T02:50:58Z")

</div>

Hello, have you solved this GPUs problem?  
I am also unable to GPUs inference Thanks for your help

---

<div class="post-metadata">

**Author:** ![curtislisle](https://avatars.discourse-cdn.com/v4/letter/c/48db29/32.png) [@curtislisle](https://discourse.slicer.org/u/curtislisle)\
**Post date:** [August 1, 2024, 3:31am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/15 "2024-08-01T03:31:05Z")

</div>

It has been a while since I used MONAILabel, but if it uses CPU only and not GPU on your system, it might be that you have the cpu version of torch installed instead of the GPU equipped version of torch. Or the Nvidia drivers are not installed correctly.

Whether a model uses multiple GPUs is generally controlled by how the inferencing python code is written, so you would need to look at the way the inferencing is performed inside the MONAILabel server to see if it is written for multi GPU. In my limited experience, it is often possible to change the code of a PyTorch model to inference across multi GPUs on a single system.

Multi-GPU works on Linux systems when the Nvidia drivers and torch-gpu versions are installed correctly. I have not personally tested multi-GPU on windows, but I assume it works as well.

Finally, deep learning, especially on 3D volumes, can be very memory intensive, whether it is run on CPU or GPU. Even with a GPU, the system memory must still be large enough to hold the data being written or read from the GPU. With small memory, it may help to down sample the volume or segment several smaller ROIs and then combine the segmentations back together later.

---

<div class="post-metadata">

**Author:** ![ece](https://avatars.discourse-cdn.com/v4/letter/e/3be4f8/32.png) [@ece](https://discourse.slicer.org/u/ece)\
**Post date:** [August 1, 2024, 4:52am UTC](https://discourse.slicer.org/t/monailabel-cuda-out-of-memory-with-bundle/29914/17 "2024-08-01T04:52:49Z")

</div>

Hello, thank you for your reply. I use a Linux system. I’m using Nvidia drivers and a torch-gpu. I want to use multi-GPU inference. I saw that you mentioned that multiple Gpus work in Linux, how to modify the implementation of thank you for your reply
