BUG: pandas 1.5 fails to groupby on (nullable) Int64 column with dropna=False · Issue #48794 · pandas-dev/pandas (original) (raw)

import numpy as np import pandas as pd

df = pd.DataFrame({ 'a':np.random.randn(150), 'b':np.repeat( [25,50,75,100,np.nan], 30), 'c':np.random.choice( ['panda','python','shark'], 150), })

df['b'] = df['b'].astype("Int64") df = df.set_index(['b', 'c'])

groups = df.groupby(axis='rows', level='b', dropna=False)

Crashes here with something like "index 91 is out of bounds for axis 0 with size 4"

groups.ngroups

Attempting to run df.groupby() with an Int64Dtype column (nullable integers) no longer works as of Pandas 1.5.0. Seems to be caused by #46601, and I'm unsure if the patch at #48702 fixes this issue.

Traceback (most recent call last):
  File "pandas15_groupby_bug.py", line 17, in <module>
    groups.ngroups
  File "venv\lib\site-packages\pandas\core\groupby\groupby.py", line 992, in __getattribute__
    return super().__getattribute__(attr)
  File "venv\lib\site-packages\pandas\core\groupby\groupby.py", line 671, in ngroups
    return self.grouper.ngroups
  File "pandas\_libs\properties.pyx", line 36, in pandas._libs.properties.CachedProperty.__get__
  File "venv\lib\site-packages\pandas\core\groupby\ops.py", line 983, in ngroups
    return len(self.result_index)
  File "pandas\_libs\properties.pyx", line 36, in pandas._libs.properties.CachedProperty.__get__
  File "venv\lib\site-packages\pandas\core\groupby\ops.py", line 994, in result_index
    return self.groupings[0].result_index.rename(self.names[0])
  File "pandas\_libs\properties.pyx", line 36, in pandas._libs.properties.CachedProperty.__get__
  File "venv\lib\site-packages\pandas\core\groupby\grouper.py", line 648, in result_index
    return self.group_index
  File "pandas\_libs\properties.pyx", line 36, in pandas._libs.properties.CachedProperty.__get__
  File "venv\lib\site-packages\pandas\core\groupby\grouper.py", line 656, in group_index
    uniques = self._codes_and_uniques[1]
  File "pandas\_libs\properties.pyx", line 36, in pandas._libs.properties.CachedProperty.__get__
  File "venv\lib\site-packages\pandas\core\groupby\grouper.py", line 693, in _codes_and_uniques
    codes, uniques = algorithms.factorize(  # type: ignore[assignment]
  File "venv\lib\site-packages\pandas\core\algorithms.py", line 792, in factorize
    codes, uniques = values.factorize(  # type: ignore[call-arg]
  File "venv\lib\site-packages\pandas\core\arrays\masked.py", line 918, in factorize
    uniques = np.insert(uniques, na_code, 0)
  File "<__array_function__ internals>", line 180, in insert
  File "venv\lib\site-packages\numpy\lib\function_base.py", line 5332, in insert
    raise IndexError(f"index {obj} is out of bounds for axis {axis} "
IndexError: index 91 is out of bounds for axis 0 with size 4

Groupby shouldn't throw, and behave the same as on Pandas 1.4.