Deploying Medical Large Language Models on Consumer-Grade Hardware: A Scoping Review of Quantization, Parameter-Efficient Fine-Tuning, and Inference Latency

Authors

  • Kevin Kipmutai Co-operative University of Kenya
  • Dancun Nyale Co-operative University of Kenya
  • Stanley Kiplagat Rotich Machakos University

Keywords:

Quantization, QLoRA, consumer GPU, medical LLM, low-resource setting

Abstract

There is a challenge of limited access to high-performance computing infrastructure, unreliable electricity, and intermittent internet connectivity in deployment of large language models (LLM) in clinical settings in low- and middle-income countries (LMICs). This scoping review aims to compile evidence regarding the use of medical LLMs on consumer-grade devices using quantization and parameter-efficient fine-tuning (PEFT). Based on the PRISMA extension for scoping reviews (PRISMA-ScR), we conducted searches across PubMed Central, medRxiv, arXiv, and Google Scholar for studies published between 2024 and 2026 on quantized or PEFT-adapted LLMs with reported hardware, memory, latency or accuracy values for medical tasks. Eight studies were found that fulfilled the inclusion criteria, covering clinical data extraction, biomedical question answering, radiology report generation, diagnosis in intensive care, medical text simplification, and diagnostic support in East Africa. Deployments on consumer GPUs like the 3-billion-parameter RTX 3060 (12 GB) and even an 8 GB laptop delivered response times below 4 seconds, while one offline diagnostic system on 3-billion-parameters achieved 100% top-3 diagnostic accuracy on a set of curated clinical cases from East Africa. To test for clinical safety, a benchmark of 26 quantized models used for intensive care showed that most of the models failed a simple patient-safety test, although they responded quickly and with competence. Computational feasibility does not measure clinical safety. Evidence from Africa and other LMIC contexts is still limited, with just one of eight studies having taken place in a sub-Saharan African context, and no studies reporting a completed prospective clinical deployment. The results indicate that quantizing and PEFT enable medical LLMs to be practically used in resource-limited settings, yet clinical validation, safety assessment, and evidence specific to LMICs are still significant challenges.

References

Al-Ganad, A., Al-Shahdhi, A., Al-Dhaifi, O., Hajeb, E., Hajeb, H., & Al-Motarreb, A. (2026). Deploying medical AI in low-resource settings: A scoping review of challenges and strategies. Frontiers in Digital Health, 8, 1743634. https://doi.org/10.3389/fdgth.2026.1743634

Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arber, S., von Arx, S., ... & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.

D'addario, A. M. V. (2025). Breaking the cost barrier: How quantization enables efficient development and deployment of LLMs for public healthcare. medRxiv. https://doi.org/10.1101/2025.11.17.25340460 (preprint, not peer-reviewed)

Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs. arXiv preprint arXiv:2305.14314.

Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2023). GPTQ: Accurate post-training quantization for generative pre-trained transformers. International Conference on Learning Representations (ICLR).

Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., ... & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations (ICLR).

Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., ... & Amodei, D. (2020). Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.

Kim, S. H., Schramm, S., Adams, L. C., Braren, R., Bressem, K. K., Keicher, M., Platzek, P.-S., Paprottka, K. J., Zimmer, C., Hedderich, D. M., & Wiestler, B. (2025). Benchmarking the diagnostic performance of open source LLMs in 1933 Eurorad case reports. npj Digital Medicine, 8, 97. https://doi.org/10.1038/s41746-025-01488-3

Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., ... & Stoica, I. (2023). Efficient memory management for large language model serving with paged attention. Proceedings of the 29th Symposium on Operating Systems Principles (SOSP), 611-626.

Lee, P., Bubeck, S., & Petro, J. (2023). Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. New England Journal of Medicine, 388(13), 1233-1239.

Lin, J., Tang, J., Tang, S., Chen, S., Yang, W., Chen, W. M., ... & Han, S. (2024). AWQ: Activation-aware weight quantization for LLM compression and acceleration. Proceedings of Machine Learning and Systems (MLSys), 6, 87-100.

Meskó, B., & Görög, M. (2020). A short guide for medical professionals in the era of artificial intelligence. NPJ Digital Medicine, 3(1), 126.

Meta AI. (2024). Llama 3 model card. GitHub. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md

Moor, M., Banerjee, O., Abad, Z. S. H., Krumholz, H. M., Leskovec, J., Topol, E. J., & Rajpurkar, P. (2023). Foundation models for generalist medical artificial intelligence. Nature, 616(7956), 259-265.

Rajbhandari, S., Rasley, J., Ruwase, O., & He, Y. (2020). ZeRO: Memory optimizations toward training trillion parameter models. Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC20), 1-16.

Shakya, P. R., Khaneja, A., & Wagholikar, K. B. (2025). For clinical data extraction, QLoRA attains accuracy close to LoRA while requiring lower compute resources. PMC12633606. https://pmc.ncbi.nlm.nih.gov/articles/PMC12633606/

Shlyakhta, T. (2026). Benchmarking large language models for intensive care unit clinical decision support: A dual safety evaluation of 26 models on consumer hardware. medRxiv. https://doi.org/10.64898/2026.02.08.26345854 (preprint, not peer-reviewed)

Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., ... & Natarajan, V. (2023). Large language models encode clinical knowledge. Nature, 620(7972), 172-180.

Srinivasu, P. N., Samudrala, A., Devesh, P. G., Reddy, S. A., & Al Nuaim, A. (2026). Quantized low-rank adaptation in large language models for clinical text simplification. International Journal of Computational Intelligence Systems, 19, 267. https://doi.org/10.1007/s44196-026-01402-z

Tricco, A. C., Lillie, E., Zarin, W., O'Brien, K. K., Colquhoun, H., Levac, D., ... & Straus, S. E. (2018). PRISMA extension for scoping reviews (PRISMA-ScR): Checklist and explanation. Annals of Internal Medicine, 169(7), 467-473. https://doi.org/10.7326/M18-0850

Tu, T., Azizi, S., Driess, D., Schaekermann, M., Amin, M., Chang, P. C., ... & Natarajan, V. (2024). Towards generalist biomedical AI. NEJM AI, 1(3), AIoa2300138.

Ugarte-Gil, C., Icochea, M., Llontop Otero, J. C., Villaizan, K., Young, N., Cao, Y., Liu, B., Griffin, T., & Brunette, M. J. (2020). Implementing a socio-technical system for computer-aided tuberculosis diagnosis in Peru: A field trial among health professionals in resource-constraint settings. Health Informatics Journal, 26(4), 2762-2775.

Valli, A. D., Tingle, S. J., Kazerouni, S., Kalpana, T. S. R., Karki, B., Knight, S. R., Wilson, C., & Kourounis, G. (2026). Artificial intelligence in surgical care within low-income and middle-income countries: A scoping review of development, validation, and deployment. eClinicalMedicine, 94, 103836. https://doi.org/10.1016/j.eclinm.2026.103836

Voinea, S.-V., Mămuleanu, M., Teică, R. V., Florescu, L. M., Selișteanu, D., & Gheonea, I. A. (2024). GPT-driven radiology report generation with fine-tuned Llama 3. Bioengineering, 11(10), 1043. https://doi.org/10.3390/bioengineering11101043

Wahl, B., Cossy-Gantner, A., Germann, S., & Schwalbe, N. R. (2018). Artificial intelligence (AI) and global health: How can AI contribute to health in resource-poor settings? BMJ Global Health, 3(4), e000798.

Walusimbi, J., Oguti, A. M., Sserwadda, A., Kasasira, P. B., & Okoboi, C. B. (2026). Aletheia: An offline-first clinical decision support system for differential diagnosis in low-resource healthcare settings. arXiv preprint arXiv:2607.24814.

World Health Organization. (2021). Ethics and governance of artificial intelligence for health: WHO guidance. World Health Organization.

Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., & Han, S. (2023). SmoothQuant: Accurate and efficient post-training quantization for large language models. Proceedings of the 40th International Conference on Machine Learning (ICML), 38087-38099.

Zhan, Z., Zhou, S., Zeng, M., Yu, K., Song, M., Chen, X., Wang, J., Hou, Y., & Zhang, R. (2025). Quantized large language models in biomedical natural language processing: Evaluation and recommendation. arXiv preprint arXiv:2509.04534.

Downloads

Published

2026-09-24

How to Cite

Kevin Kipmutai, Dancun Nyale, & Stanley Kiplagat Rotich. (2026). Deploying Medical Large Language Models on Consumer-Grade Hardware: A Scoping Review of Quantization, Parameter-Efficient Fine-Tuning, and Inference Latency. African Journal of Education,Science and Technology (AJEST), 8(4), 243–254. Retrieved from https://ajest.org/index.php/ajest/article/view/1037

Issue

Section

Articles

Similar Articles

<< < 5 6 7 8 9 10 11 12 13 14 > >> 

You may also start an advanced similarity search for this article.