# Getting started with MONAIlabel

**URL:** <https://discourse.slicer.org/t/getting-started-with-monailabel/38559>\
**Category:** Support\
**Tags:** monai, monailabel\
**Created:** [September 26, 2024, 1:01pm UTC](https://discourse.slicer.org/t/getting-started-with-monailabel/38559 "2024-09-26T13:01:05Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![gendex](https://avatars.discourse-cdn.com/v4/letter/g/ea5d25/32.png) [@gendex](https://discourse.slicer.org/u/gendex)\
**Post date:** [September 26, 2024, 1:01pm UTC](https://discourse.slicer.org/t/getting-started-with-monailabel/38559/1 "2024-09-26T13:01:05Z")

</div>

Dear Slicer Community,

I have been working with Slicer for a while now and have recently started using the MONAI module in Slicer. However, I’m having some issues getting the output that I’m looking for.

I’ve managed to run the segmentation model to train an initial dataset that I have previously segmented using TotalSegmentator. However, when testing the model there is a large offset in the output as seen below.

 ![output_offset](https://us1.discourse-cdn.com/flex002/uploads/slicer/original/3X/2/a/2a608243f9f1d6a1021755e7db2fa9536b0f01e5.png)

I tried to recreate this by training the segmentation model on a subset of the publicly available TotalSegmentator dataset and encountered different errors. It seems that MONAI is struggling with applying a transform (see below):

_[2024-09-24 12:20:25,830] [17576] [MainThread] [INFO] (ignite.engine.engine.SupervisedEvaluator:876) - Engine run resuming from iteration 0, epoch 0 until 1 epochs_  
_[2024-09-24 12:32:03,933] [17576] [MainThread] [ERROR] (ignite.engine.engine.SupervisedEvaluator:1086) - Current run is terminating due to exception: applying transform \<monai.transforms.compose.Compose object at 0x000001D6A6078CD0\>_  
_2024-09-24 12:32:03,933 - ERROR - Exception: applying transform \<monai.transforms.compose.Compose object at 0x000001D6A6078CD0\>_

_…_

_torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 3.16 GiB. GPU_

_…_

_RuntimeError: applying transform \<monai.transforms.post.dictionary.AsDiscreted object at 0x000001D6A60781C0\>_

What kind of transform is MONAI trying to perform?  
Could it be a driver issue that CUDA is running out of memory at 3.16GiB? The used GPU is an Nvidia RTX 4090 (24GiB).  
Do I need to manually clear the GPU cache?

When I tried to train a different subset of my own data afterwards it would take about a minute after every epoch to resume the engine run (see below).

_[2024-09-24 13:45:26,169] [16008] [MainThread] [INFO] (ignite.engine.engine.SupervisedEvaluator:876) - Engine run resuming from iteration 0, epoch 0 until 1 epochs_  
_[2024-09-24 13:46:12,469] [16008] [MainThread] [INFO] (ignite.engine.engine.SupervisedEvaluator:259) - Got new best metric of val\_mean\_dice: 0.0_  
_2024-09-24 13:46:12,469 - INFO - Epoch[1] Metrics – val\_mean\_dice: 0.0000 val\_skeletal\_muscle\_mean\_dice: 0.0000 val\_subcutaneous\_fat\_mean\_dice: 0.0000 val\_torso\_fat\_mean\_dice: 0.0000_

During the first training that generated the offset, training was a lot quicker. Eventually, the offset was recreated regardless of training and testing both only left shoulder scans and also using mixed unilateral scans.

Finally, is there a way to compare the outputs of different models? Auto Segmentation/Models only shows “segmentation” as an option, even though the different models (all producing offset) have been saved with unique names. Could I even be testing on a wrong model?

I would highly appreciate any inputs. Thank you!

Best regards,  
Dennis

---

<div class="post-metadata">

**Author:** ![gendex](https://avatars.discourse-cdn.com/v4/letter/g/ea5d25/32.png) [@gendex](https://discourse.slicer.org/u/gendex)\
**Post date:** [October 3, 2024, 2:19pm UTC](https://discourse.slicer.org/t/getting-started-with-monailabel/38559/2 "2024-10-03T14:19:10Z")

</div>

Hello again, I’d like to provide an update on the situation, give some more background info, reorganize the issues that I’ve encountered and pose more concise questions.

I have recently started using MONAI for a research project that involves automating the segmentation of the scapula to extract bone density values that then could potentially be used for surgical planning. My background is in biomechanics with some but limited programming skills. However, I have been using 3D Slicer for almost a year now.

I have a dataset of shoulder CT scans that are cropped in a way that the crop axis is not aligned with the scanning axis and when loading them into 3D Slicer, some scans appear rotated. I used TotalSegmentator to segment the scapula and other structures of interest, and the different orientations of the scans did not seem to be a problem.

Since MONAI seems to be preprocessing the data with transforms etc., I assume that the differently oriented scans should not cause any issues when training a segmentation model in MONAI. Is that correct?

To get to know how MONAI works, I tested segmenting the spleen dataset on the pretrained spleen model and that worked fine.

As a first test using my own data, I trained a subset of 7 scans (average dimensions: 1000x500x150, average spacing: 0.25x0.25x1mm) with 3 segments each (adjusted segmentation.py in configs accordingly) using the segmentation model for 500 epochs on an Nvidia RTX 4090 with 24GiB. The training was successfully completed in about an hour.

Does this training duration fall into the expected time range?

I then wanted to test the performance of the trained model using the Auto Segmentation tab. The dropdown menu in that tab had “segmentation” as the only option. I loaded the next sample that didn’t have any ground truth segmentations and hit run. The resulting segmentations had roughly the desired shape, but they were offset so much, that they lay outside the body (see image below).

Am I testing on my own model?

How can I specify which model I want to test on?

Are these results to be expected or is that sufficient training to expect better results?

Do I need to increase the number of scans and epochs to get better results?

 ![output_offset](https://us1.discourse-cdn.com/flex002/uploads/slicer/original/3X/2/a/2a608243f9f1d6a1021755e7db2fa9536b0f01e5.png)

After restarting the server in the same app with a different dataset loaded, I trained another segmentation model giving it a different name in the options tab. After successful training, there still is only the same single model (“segmentation”) available in the Auto Segmentation tab.

Which model am I testing on if I load the next unlabeled scan and hit run?

How can I choose which model I want to test on?

Is it even possible to have multiple segmentation models in the same app or do I have to create a new app for every model?

Is there a way to visualize the segmentation outputs from the validation scans when val\_split is \>0 / when there are scans used for validation?

To check if there is something wrong with my dataset I tried to reproduce the offset results. I used a subset (16 chest CT scans) of the TotalSegmentator dataset and ran TotalSegmentator on those scans to get the needed segmentations. Then, I loaded the data into the same app and started training, again with the segmentation model. After the first epoch, it timed out giving the following error message:

_[2024-09-24 12:20:11,585] [17576] [MainThread] [INFO] (monailabel.tasks.train.basic\_train:264) - 0 - Records for Training: 16_  
_[2024-09-24 12:20:11,588] [17576] [MainThread] [INFO] (ignite.engine.engine.SupervisedTrainer:876) - Engine run resuming from iteration 0, epoch 0 until 100 epochs_  
_2024-09-24 12:20:13,821 - INFO - Epoch: 1/100, Iter: 1/16 – train\_loss: 2.7914_  
_2024-09-24 12:20:14,936 - INFO - Epoch: 1/100, Iter: 2/16 – train\_loss: 2.2761_  
_2024-09-24 12:20:15,729 - INFO - Epoch: 1/100, Iter: 3/16 – train\_loss: 1.9255_  
_2024-09-24 12:20:16,620 - INFO - Epoch: 1/100, Iter: 4/16 – train\_loss: 1.9023_  
_2024-09-24 12:20:17,308 - INFO - Epoch: 1/100, Iter: 5/16 – train\_loss: 1.9119_  
_2024-09-24 12:20:18,056 - INFO - Epoch: 1/100, Iter: 6/16 – train\_loss: 2.3988_  
_2024-09-24 12:20:18,838 - INFO - Epoch: 1/100, Iter: 7/16 – train\_loss: 1.9958_  
_2024-09-24 12:20:19,639 - INFO - Epoch: 1/100, Iter: 8/16 – train\_loss: 1.9923_  
_2024-09-24 12:20:20,427 - INFO - Epoch: 1/100, Iter: 9/16 – train\_loss: 2.3822_  
_2024-09-24 12:20:21,259 - INFO - Epoch: 1/100, Iter: 10/16 – train\_loss: 2.1640_  
_2024-09-24 12:20:22,024 - INFO - Epoch: 1/100, Iter: 11/16 – train\_loss: 1.8334_  
_2024-09-24 12:20:22,963 - INFO - Epoch: 1/100, Iter: 12/16 – train\_loss: 1.8153_  
_2024-09-24 12:20:23,869 - INFO - Epoch: 1/100, Iter: 13/16 – train\_loss: 1.9521_  
_2024-09-24 12:20:24,715 - INFO - Epoch: 1/100, Iter: 14/16 – train\_loss: 1.7093_  
_2024-09-24 12:20:25,162 - INFO - Epoch: 1/100, Iter: 15/16 – train\_loss: 2.0465_  
_2024-09-24 12:20:25,813 - INFO - Epoch: 1/100, Iter: 16/16 – train\_loss: 2.4063_  
_[2024-09-24 12:20:25,819] [17576] [MainThread] [INFO] (ignite.engine.engine.SupervisedTrainer:259) - Got new best metric of train\_mean\_dice: 0.04801511764526367_  
_2024-09-24 12:20:25,819 - INFO - Epoch[1] Metrics – train\_mean\_dice: 0.0480 train\_skeletal\_muscle\_mean\_dice: 0.1105 train\_subcutaneous\_fat\_mean\_dice: 0.0133 train\_torso\_fat\_mean\_dice: 0.0000_  
_2024-09-24 12:20:25,819 - INFO - Key metric: train\_mean\_dice best value: 0.04801511764526367 at epoch: 1_  
_[2024-09-24 12:20:25,830] [17576] [MainThread] [INFO] (ignite.engine.engine.SupervisedEvaluator:876) - Engine run resuming from iteration 0, epoch 0 until 1 epochs_  
_[2024-09-24 12:32:03,933] [17576] [MainThread] [ERROR] (ignite.engine.engine.SupervisedEvaluator:1086) - Current run is terminating due to exception: applying transform \<monai.transforms.compose.Compose object at 0x000001D6A6078CD0\>_  
_2024-09-24 12:32:03,933 - ERROR - Exception: applying transform \<monai.transforms.compose.Compose object at 0x000001D6A6078CD0\>_

I don’t understand this error.

Is MONAI struggling to transform the scans?

Or is the GPU struggling to process that amount of data?

How much data can a single Nvidia RTX 4090 (24GiB) handle?

Is there a buildup of cache when training multiple models in the same app?

With this comparison failed, I set out to train another model with a different subset of my own data. This time I loaded 13 scans with roughly the same dimensions and spacing (as the original dataset) and trained again. Training took a lot longer this time. After every epoch training seemed to pause for about one minute at this point (same as where the error happened previously):

_[2024-09-24 13:45:24,528] [16008] [MainThread] [INFO] (monailabel.tasks.train.basic\_train:264) - 0 - Records for Training: 13_  
_[2024-09-24 13:45:24,530] [16008] [MainThread] [INFO] (ignite.engine.engine.SupervisedTrainer:876) - Engine run resuming from iteration 0, epoch 0 until 100 epochs_  
_2024-09-24 13:45:24,786 - INFO - Epoch: 1/100, Iter: 1/13 – train\_loss: 2.4122_  
_2024-09-24 13:45:24,878 - INFO - Epoch: 1/100, Iter: 2/13 – train\_loss: 2.3242_  
_2024-09-24 13:45:24,992 - INFO - Epoch: 1/100, Iter: 3/13 – train\_loss: 2.0153_  
_2024-09-24 13:45:25,109 - INFO - Epoch: 1/100, Iter: 4/13 – train\_loss: 1.6996_  
_2024-09-24 13:45:25,225 - INFO - Epoch: 1/100, Iter: 5/13 – train\_loss: 2.2290_  
_2024-09-24 13:45:25,340 - INFO - Epoch: 1/100, Iter: 6/13 – train\_loss: 1.4649_  
_2024-09-24 13:45:25,456 - INFO - Epoch: 1/100, Iter: 7/13 – train\_loss: 1.4484_  
_2024-09-24 13:45:25,570 - INFO - Epoch: 1/100, Iter: 8/13 – train\_loss: 1.4148_  
_2024-09-24 13:45:25,686 - INFO - Epoch: 1/100, Iter: 9/13 – train\_loss: 1.3817_  
_2024-09-24 13:45:25,802 - INFO - Epoch: 1/100, Iter: 10/13 – train\_loss: 2.7751_  
_2024-09-24 13:45:25,925 - INFO - Epoch: 1/100, Iter: 11/13 – train\_loss: 2.1342_  
_2024-09-24 13:45:26,040 - INFO - Epoch: 1/100, Iter: 12/13 – train\_loss: 1.7436_  
_2024-09-24 13:45:26,153 - INFO - Epoch: 1/100, Iter: 13/13 – train\_loss: 1.3803_  
_[2024-09-24 13:45:26,157] [16008] [MainThread] [INFO] (ignite.engine.engine.SupervisedTrainer:259) - Got new best metric of train\_mean\_dice: 0.0_  
_2024-09-24 13:45:26,157 - INFO - Epoch[1] Metrics – train\_mean\_dice: 0.0000 train\_skeletal\_muscle\_mean\_dice: 0.0000 train\_subcutaneous\_fat\_mean\_dice: 0.0000 train\_torso\_fat\_mean\_dice: 0.0000_  
_2024-09-24 13:45:26,157 - INFO - Key metric: train\_mean\_dice best value: 0.0 at epoch: 1_  
_[2024-09-24 13:45:26,169] [16008] [MainThread] [INFO] (ignite.engine.engine.SupervisedEvaluator:876) - Engine run resuming from iteration 0, epoch 0 until 1 epochs_  
_[2024-09-24 13:46:12,469] [16008] [MainThread] [INFO] (ignite.engine.engine.SupervisedEvaluator:259) - Got new best metric of val\_mean\_dice: 0.0_  
_2024-09-24 13:46:12,469 - INFO - Epoch[1] Metrics – val\_mean\_dice: 0.0000 val\_skeletal\_muscle\_mean\_dice: 0.0000 val\_subcutaneous\_fat\_mean\_dice: 0.0000 val\_torso\_fat\_mean\_dice: 0.0000_  
_2024-09-24 13:46:12,469 - INFO - Key metric: val\_mean\_dice best value: 0.0 at epoch: 1_  
_[2024-09-24 13:46:12,553] [16008] [MainThread] [INFO] (monailabel.tasks.train.handler:86) - New Model published: C:\Users\Eva\radiologyBoneDensity\model\segmentation\test\_rem\_pat\model.pt =\> C:\Users\Eva\radiologyBoneDensity\model\segmentation.pt_  
_[2024-09-24 13:46:12,554] [16008] [MainThread] [INFO] (ignite.engine.engine.SupervisedEvaluator:972) - Epoch[1] Complete. Time taken: 00:00:46.368_  
_[2024-09-24 13:46:12,555] [16008] [MainThread] [INFO] (ignite.engine.engine.SupervisedEvaluator:988) - Engine run complete. Time taken: 00:00:46.386_  
_[2024-09-24 13:46:12,607] [16008] [MainThread] [INFO] (ignite.engine.engine.SupervisedTrainer:972) - Epoch[1] Complete. Time taken: 00:00:48.035_  
_2024-09-24 13:46:12,835 - INFO - Epoch: 2/100, Iter: 1/13 – train\_loss: 1.7215_  
_2024-09-24 13:46:13,029 - INFO - Epoch: 2/100, Iter: 2/13 – train\_loss: 1.3821_  
_2024-09-24 13:46:13,217 - INFO - Epoch: 2/100, Iter: 3/13 – train\_loss: 1.6698_  
_2024-09-24 13:46:13,416 - INFO - Epoch: 2/100, Iter: 4/13 – train\_loss: 1.7367_  
_2024-09-24 13:46:13,609 - INFO - Epoch: 2/100, Iter: 5/13 – train\_loss: 1.7043_  
_2024-09-24 13:46:13,824 - INFO - Epoch: 2/100, Iter: 6/13 – train\_loss: 1.6097_  
_2024-09-24 13:46:14,019 - INFO - Epoch: 2/100, Iter: 7/13 – train\_loss: 2.4585_  
_2024-09-24 13:46:14,211 - INFO - Epoch: 2/100, Iter: 8/13 – train\_loss: 1.6047_  
_2024-09-24 13:46:14,419 - INFO - Epoch: 2/100, Iter: 9/13 – train\_loss: 1.8360_  
_2024-09-24 13:46:14,642 - INFO - Epoch: 2/100, Iter: 10/13 – train\_loss: 1.3524_  
_2024-09-24 13:46:14,860 - INFO - Epoch: 2/100, Iter: 11/13 – train\_loss: 1.7037_  
_2024-09-24 13:46:15,045 - INFO - Epoch: 2/100, Iter: 12/13 – train\_loss: 1.4732_  
_2024-09-24 13:46:15,241 - INFO - Epoch: 2/100, Iter: 13/13 – train\_loss: 1.3968_  
_[2024-09-24 13:46:15,256] [16008] [MainThread] [INFO] (ignite.engine.engine.SupervisedTrainer:259) - Got new best metric of train\_mean\_dice: 0.04763128608465195_  
_2024-09-24 13:46:15,257 - INFO - Epoch[2] Metrics – train\_mean\_dice: 0.0476 train\_skeletal\_muscle\_mean\_dice: 0.0953 train\_subcutaneous\_fat\_mean\_dice: 0.0000 train\_torso\_fat\_mean\_dice: 0.0000_  
_2024-09-24 13:46:15,257 - INFO - Key metric: train\_mean\_dice best value: 0.04763128608465195 at epoch: 2_  
_[2024-09-24 13:46:15,259] [16008] [MainThread] [INFO] (ignite.engine.engine.SupervisedEvaluator:876) - Engine run resuming from iteration 0, epoch 1 until 2 epochs_  
_[2024-09-24 13:47:01,627] [16008] [MainThread] [INFO] (ignite.engine.engine.SupervisedEvaluator:259) - Got new best metric of val\_mean\_dice: 0.024709107354283333_  
_2024-09-24 13:47:01,627 - INFO - Epoch[2] Metrics – val\_mean\_dice: 0.0247 val\_skeletal\_muscle\_mean\_dice: 0.0741 val\_subcutaneous\_fat\_mean\_dice: 0.0000 val\_torso\_fat\_mean\_dice: 0.0000_

I expected training to take longer in a somewhat exponential way, but I didn’t expect such a big difference (around 5secs per epoch in the first training with 7 samples vs 1min per epoch with 13 samples).

What is happening in that step?

Is this increase in training time expected?

Any input would be greatly appreciated.

---

<div class="post-metadata">

**Author:** ![pieper](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.slicer.org/pieper/32/8_2.png) [@pieper](https://discourse.slicer.org/u/pieper)\
**Post date:** [October 4, 2024, 1:56pm UTC](https://discourse.slicer.org/t/getting-started-with-monailabel/38559/3 "2024-10-04T13:56:36Z")

</div>

> [@gendex](#):
>
> Since MONAI seems to be preprocessing the data with transforms etc., I assume that the differently oriented scans should not cause any issues when training a segmentation model in MONAI. Is that correct?

I would guess not and that you need to resample the volumes and labelmaps to the same pixel space for MONAI.

As for your other questions, they may get resolved after you get past this first issue. But if not, maybe try asking one small question per post.
