In a multipe gpu setup, https://github.com/ModelCloud/GPTQModel/pull/39 may need to execute on a single gpu in a multi-gpu env where the gpu_id is not 0 according to nvidia-smi. The current setup assume TVM target is always the first gpu (index=0) which is incorrect. Fix = pass in gpu_id: int and make nvidia-smi --id use it
Nvidia makes several oem, non-public, version of A100 such as PG506-230 which is part of the official and opensource driver supported device list. They are essentially A100 with different VRAM sizes. Remap this model so TVM can be matched to A100 correctly. There may be other gpus affected so added TODO to move this into a helper re-mapper method. Fix = manual remap.
TEST
[x] PASSED on PG506-230 => remap to A100
[x] PASSED on 4090 in a multi-gpu setup where id is > 0
There are two bug fixes in this PR
In a multipe gpu setup, https://github.com/ModelCloud/GPTQModel/pull/39 may need to execute on a single gpu in a multi-gpu env where the gpu_id is not 0 according to nvidia-smi. The current setup assume TVM target is always the first gpu (index=0) which is incorrect. Fix = pass in
gpu_id: int
and makenvidia-smi --id
use itNvidia makes several oem, non-public, version of A100 such as PG506-230 which is part of the official and opensource driver supported device list. They are essentially A100 with different VRAM sizes. Remap this model so TVM can be matched to A100 correctly. There may be other gpus affected so added TODO to move this into a helper re-mapper method. Fix = manual remap.
TEST
@LeiWang1999