Finetuning NeMo parakeet in google colab results CUDA_ERROR_UNSUPPORTED_PTX_VERSION
12:05 26 May 2025

Aim: I want to finetune parakeet v2 model to a different dataset. I picked LJ dataset just to make myself familiar with the finetuning process.

For doing this I ran the following notebook

This works fine on kaggle while on google colab it gives me the following error.

File "/usr/local/lib/python3.11/dist-packages/nemo/collections/asr/parts/numba/rnnt_loss/utils/cuda_utils/gpu_rnnt.py", line 768, in cost_and_grad
    return self.compute_cost_and_score(
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.11/dist-packages/nemo/collections/asr/parts/numba/rnnt_loss/utils/cuda_utils/gpu_rnnt.py", line 604, in compute_cost_and_score
    self.log_softmax(label_acts, denom)
  File "/usr/local/lib/python3.11/dist-packages/nemo/collections/asr/parts/numba/rnnt_loss/utils/cuda_utils/gpu_rnnt.py", line 107, in log_softmax
    reduce.reduce_max(
  File "/usr/local/lib/python3.11/dist-packages/nemo/collections/asr/parts/numba/rnnt_loss/utils/cuda_utils/reduce.py", line 353, in reduce_max
    return ReduceHelper(
           ^^^^^^^^^^^^^
  File "/usr/local/lib/python3.11/dist-packages/nemo/collections/asr/parts/numba/rnnt_loss/utils/cuda_utils/reduce.py", line 294, in ReduceHelper
    _reduce_rows[grid_size, CTA_REDUCE_SIZE, stream, 0](I_opid, R_opid, acts, output, num_rows)
  File "/usr/local/lib/python3.11/dist-packages/numba_cuda/numba/cuda/dispatcher.py", line 608, in __call__
    return self.dispatcher.call(args, self.griddim, self.blockdim,
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.11/dist-packages/numba_cuda/numba/cuda/dispatcher.py", line 750, in call
    kernel = _dispatcher.Dispatcher._cuda_call(self, *args)
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.11/dist-packages/numba_cuda/numba/cuda/dispatcher.py", line 758, in _compile_for_args
    return self.compile(tuple(argtypes))
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.11/dist-packages/numba_cuda/numba/cuda/dispatcher.py", line 1005, in compile
    kernel.bind()
  File "/usr/local/lib/python3.11/dist-packages/numba_cuda/numba/cuda/dispatcher.py", line 257, in bind
    cufunc = self._codelibrary.get_cufunc()
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.11/dist-packages/numba_cuda/numba/cuda/codegen.py", line 248, in get_cufunc
    cubin = self.get_cubin(cc=device.compute_capability)
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.11/dist-packages/numba_cuda/numba/cuda/codegen.py", line 227, in get_cubin
    self._link_all(linker, cc, ignore_nonlto=False)
  File "/usr/local/lib/python3.11/dist-packages/numba_cuda/numba/cuda/codegen.py", line 189, in _link_all
    linker.add_ptx(ptx.encode())
  File "/usr/local/lib/python3.11/dist-packages/numba_cuda/numba/cuda/cudadrv/driver.py", line 2904, in add_ptx
    raise LinkerError("%s\n%s" % (e, self.error_log))
numba.cuda.cudadrv.driver.LinkerError: [222] Call to cuLinkAddData results in CUDA_ERROR_UNSUPPORTED_PTX_VERSION
ptxas application ptx input, line 9; fatal   : Unsupported .version 8.5; current version is '8.4'

So I tried to reproduce a minimal version of the error

import os
from numba import cuda
import numpy as np
@cuda.jit
def test_kernel(arr):
    i = cuda.grid(1)
    if i < arr.size:
        arr[i] += 1

arr = np.zeros(100, dtype=np.float32)
d_arr = cuda.to_device(arr)
test_kernel[100, 1](d_arr)
print(d_arr.copy_to_host())

I get the same error as above on colab.

Now following the discussion here I was successful in removing this error using

!uv pip install -q --system numba-cuda==0.4.0 --force-reinstall
!uv pip install -q --system numpy==1.26.4 --force-reinstall 

and then explicitly setting

from numba import config
config.CUDA_ENABLE_PYNVJITLINK = 1
import os
os.environ["NUMBA_CUDA_DEFAULT_PTX_CC"] = "7.5"

Now I ran the small code snippet with the these flag and my error is gone. Someone is modifying the code here

But We need to set this flag for NeMo for the WHOLE library.

I am not sure how to do this? One way is to set here but Do I Have to set everywhere?

Is there a way to set globally in every module using numba library?

cuda google-colaboratory speech-to-text numba fine-tuning