Enterprise LLM Router: Learning Quality–Capacity–Capability Trade-offs from 2026 Model Metadata

Authors

  • Lily Peng Computer Science, University Of Leeds, Leeds, WYS, UK

DOI:

https://doi.org/10.61424/jcsit.v3i1.953

Keywords:

Large language models; model routing; enterprise AI; multi-criteria selection; benchmark metadata; rate limits; capability constraints; provider generalization

Abstract

Enterprise language-model selection is a constrained decision problem, not a single leaderboard lookup. This study integrated three 2026 metadata tables covering 22 models from eight providers and evaluated benchmark quality, capability requirements, paid rate-limit capacity, and discount support. The files joined exactly on provider and model without missing cells or duplicate keys. Quality was the mean percentile rank of MMLU, HumanEval, MATH, and Arena Elo; capacity was the mean paid-RPM and paid-TPM percentile for 19 models with numeric limits; discount support averaged batch and cached-input discounts. Nested leave-one-model-out and leave-one-provider-out tests compared ridge regression, k-nearest neighbors, random forest, and a training-mean baseline. The router was evaluated on 20 capability-gated candidate sets under five utility regimes, creating 100 scenarios, followed by a 0.1-resolution weight analysis. Ridge achieved leave-one-model-out MAE 13.09 and Spearman ρ .771, but provider-held-out performance weakened to MAE 22.35 and R² −.082. Capability breadth produced the lowest routing regret (5.60), ahead of ridge (8.46) and observed-quality routing (10.60). The eight-model Pareto frontier exposed specialization: o1 led quality, Gemini 2.5 Flash led numeric paid capacity, and DeepSeek V3 combined high capacity with maximum discount support. The evidence favors transparent capability filtering and abstention when price, latency, deployment license, or context evidence is absent.

References

Aggarwal, P., Madaan, A., Anand, A., Potharaju, S. P., Mishra, S., Zhou, P., Gupta, A., Rajagopal, D., Kappaganthu, K., Yang, Y., Upadhyay, S., Faruqui, M., & Mausam. (2024). AutoMix: Automatically mixing language models. In Advances in Neural Information Processing Systems (Vol. 37). https://arxiv.org/abs/2310.12963

Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B., Tumanov, A., & Ramjee, R. (2024). Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve. In 18th USENIX Symposium on Operating Systems Design and Implementation (pp. 117–134). https://www.usenix.org/conference/osdi24/presentation/agrawal

Anthropic. (n.d.). Rate limits. Retrieved July 17, 2026, from https://platform.claude.com/docs/en/api/rate-limits

Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., Dong, Y., Tang, J., & Li, J. (2024). LongBench: A bilingual, multitask benchmark for long context understanding. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (pp. 3119–3137). https://doi.org/10.18653/v1/2024.acl-long.172

Chen, L., Zaharia, M., & Zou, J. (2020). FrugalML: How to use ML prediction APIs more accurately and cheaply. In Advances in Neural Information Processing Systems (Vol. 33, pp. 10685–10696). https://arxiv.org/abs/2006.07512

Chen, L., Zaharia, M., & Zou, J. (2024). FrugalGPT: How to use large language models while reducing cost and improving performance. Transactions on Machine Learning Research. https://openreview.net/forum?id=cSimKw5p6R

Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. de O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., . . . Zaremba, W. (2021). Evaluating large language models trained on code [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2107.03374

Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M. I., Gonzalez, J. E., & Stoica, I. (2024). Chatbot Arena: An open platform for evaluating LLMs by human preference. In Proceedings of the 41st International Conference on Machine Learning (pp. 8359–8388). https://proceedings.mlr.press/v235/chiang24b.html

Ding, D., Mallick, A., Wang, C., Sim, R., Mukherjee, S., Ruhle, V., Lakshmanan, L. V. S., & Awadallah, A. H. (2024). Hybrid LLM: Cost-efficient and quality-aware query routing [Preprint]. arXiv. https://arxiv.org/abs/2404.14618

Google. (n.d.). Rate limits. Retrieved July 17, 2026, from https://ai.google.dev/gemini-api/docs/rate-limits

Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., & Steinhardt, J. (2021a). Measuring massive multitask language understanding. In International Conference on Learning Representations. https://openreview.net/forum?id=d7KBjmI3GmQ

Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., & Steinhardt, J. (2021b). Measuring mathematical problem solving with the MATH dataset. In Advances in Neural Information Processing Systems (Vol. 34, pp. 3057–3074). https://proceedings.neurips.cc/paper/2021/hash/be83ab3ecd0db773eb2dc1b0a17836a1-Abstract.html

Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., Zhang, Y., & Ginsburg, B. (2024). RULER: What’s the real context size of your long-context language models? In First Conference on Language Modeling. https://arxiv.org/abs/2404.06654

Hu, Q. J., Bieker, J., Li, X., Jiang, N., Keigwin, B., Ranganath, G., Keutzer, K., & Upadhyay, S. K. (2024). RouterBench: A benchmark for multi-LLM routing system [Preprint]. arXiv. https://arxiv.org/abs/2403.12031

Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., & Stoica, I. (2023). Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles. https://doi.org/10.1145/3600006.3613165

Li, C., Liu, G., & Zhao, Z. (2026). Cost-aware LLM-style routing for AIOps log analysis: Log parsing, anomaly detection, fault diagnosis, and incident summarization on LogEval task files. Journal of Technology Informatics and Engineering, 5(2), 91–103. https://doi.org/10.51903/jtie.v5i2.538

Li, H., Zhang, Y., Guo, Z., Wang, C., Tang, S., Zhang, Q., Chen, Y., Qi, B., Ye, P., Bai, L., Wang, Z., & Hu, S. (2026). LLMRouterBench: A unified benchmark for robust and efficient LLM routing [Preprint]. arXiv. https://arxiv.org/abs/2601.07206

Lu, K., Yuan, H., Lin, R., Lin, J., Yuan, Z., Zhou, C., & Zhou, J. (2024). Routing to the expert: Efficient reward-guided ensemble of large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 1964–1974). https://doi.org/10.18653/v1/2024.naacl-long.109

Mohammadshahi, A., Shaikh, A. R., & Yazdani, M. (2024). Routoo: Learning to route to large language models effectively [Preprint]. arXiv. https://arxiv.org/abs/2401.13979

Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., & Stoica, I. (2025). RouteLLM: Learning to route LLMs from preference data. In International Conference on Learning Representations. https://openreview.net/forum?id=8sSqNntaMr

OpenAI. (n.d.). Rate limits. Retrieved July 17, 2026, from https://developers.openai.com/api/docs/guides/rate-limits

Shah, S., & Shridhar, K. (2025). Select-then-route: A two-stage framework for efficient LLM routing. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track (pp. 425–441). https://doi.org/10.18653/v1/2025.emnlp-industry.28

Shen, Y., Liu, Y., Huang, Z., Yin, R., Zheng, X., & Huang, X. (2025). SATER: Cost-aware and latency-efficient LLM routing through speculative answer generation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 10515–10529). https://doi.org/10.18653/v1/2025.emnlp-main.531

Shnitzer, T., Ou, A., Silva, M., Soule, K., Sun, Y., Solomon, J., Thompson, N., & Yurochkin, M. (2023). Large language model routing with benchmark datasets [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2309.15789

Uddin, I., & Bauer, A. (2026). Conformal LLM routing with distribution-free safety guarantees. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop (pp. 791–799). https://doi.org/10.18653/v1/2026.acl-srw.70

Varangot-Reille, C., Bouvard, C., Gourru, A., Ciancone, M., Schaeffer, M., & Jacquenet, F. (2025). Doing more with less: A survey on routing strategies for resource optimisation in large language model-based systems [Preprint]. arXiv. https://arxiv.org/abs/2502.00409

Wang, X., Liu, Y., Cheng, W., Zhao, X., Chen, Z., Yu, W., Fu, Y., & Chen, H. (2025). MixLLM: Dynamic routing in mixed large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 10912–10922). https://doi.org/10.18653/v1/2025.naacl-long.545

Xin, Q. (2025). Hybrid cloud architecture for efficient and cost-effective large language model deployment. Journal of Information Systems and Informatics, 7(3), 2182–2195. https://doi.org/10.51519/journalisi.v7i3.1170

Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena [Preprint]. arXiv. https://arxiv.org/abs/2306.05685

Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., & Zhang, H. (2024). DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. In 18th USENIX Symposium on Operating Systems Design and Implementation (pp. 193–210). https://www.usenix.org/conference/osdi24/presentation/zhong-yinmin

Downloads

Published

2026-07-19

How to Cite

Peng, L. (2026). Enterprise LLM Router: Learning Quality–Capacity–Capability Trade-offs from 2026 Model Metadata. Journal of Computer Science and Information Technology, 3(1), 104–120. https://doi.org/10.61424/jcsit.v3i1.953