Enterprise LLM Router: Learning Quality–Capacity–Capability Trade-offs from 2026 Model Metadata
DOI:
https://doi.org/10.61424/jcsit.v3i1.953Keywords:
Large language models; model routing; enterprise AI; multi-criteria selection; benchmark metadata; rate limits; capability constraints; provider generalizationAbstract
Enterprise language-model selection is a constrained decision problem, not a single leaderboard lookup. This study integrated three 2026 metadata tables covering 22 models from eight providers and evaluated benchmark quality, capability requirements, paid rate-limit capacity, and discount support. The files joined exactly on provider and model without missing cells or duplicate keys. Quality was the mean percentile rank of MMLU, HumanEval, MATH, and Arena Elo; capacity was the mean paid-RPM and paid-TPM percentile for 19 models with numeric limits; discount support averaged batch and cached-input discounts. Nested leave-one-model-out and leave-one-provider-out tests compared ridge regression, k-nearest neighbors, random forest, and a training-mean baseline. The router was evaluated on 20 capability-gated candidate sets under five utility regimes, creating 100 scenarios, followed by a 0.1-resolution weight analysis. Ridge achieved leave-one-model-out MAE 13.09 and Spearman ρ .771, but provider-held-out performance weakened to MAE 22.35 and R² −.082. Capability breadth produced the lowest routing regret (5.60), ahead of ridge (8.46) and observed-quality routing (10.60). The eight-model Pareto frontier exposed specialization: o1 led quality, Gemini 2.5 Flash led numeric paid capacity, and DeepSeek V3 combined high capacity with maximum discount support. The evidence favors transparent capability filtering and abstention when price, latency, deployment license, or context evidence is absent.
References
Aggarwal, P., Madaan, A., Anand, A., Potharaju, S. P., Mishra, S., Zhou, P., Gupta, A., Rajagopal, D., Kappaganthu, K., Yang, Y., Upadhyay, S., Faruqui, M., & Mausam. (2024). AutoMix: Automatically mixing language models. In Advances in Neural Information Processing Systems (Vol. 37). https://arxiv.org/abs/2310.12963
Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B., Tumanov, A., & Ramjee, R. (2024). Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve. In 18th USENIX Symposium on Operating Systems Design and Implementation (pp. 117–134). https://www.usenix.org/conference/osdi24/presentation/agrawal
Anthropic. (n.d.). Rate limits. Retrieved July 17, 2026, from https://platform.claude.com/docs/en/api/rate-limits
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., Dong, Y., Tang, J., & Li, J. (2024). LongBench: A bilingual, multitask benchmark for long context understanding. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (pp. 3119–3137). https://doi.org/10.18653/v1/2024.acl-long.172
Chen, L., Zaharia, M., & Zou, J. (2020). FrugalML: How to use ML prediction APIs more accurately and cheaply. In Advances in Neural Information Processing Systems (Vol. 33, pp. 10685–10696). https://arxiv.org/abs/2006.07512
Chen, L., Zaharia, M., & Zou, J. (2024). FrugalGPT: How to use large language models while reducing cost and improving performance. Transactions on Machine Learning Research. https://openreview.net/forum?id=cSimKw5p6R
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. de O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., . . . Zaremba, W. (2021). Evaluating large language models trained on code [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2107.03374
Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M. I., Gonzalez, J. E., & Stoica, I. (2024). Chatbot Arena: An open platform for evaluating LLMs by human preference. In Proceedings of the 41st International Conference on Machine Learning (pp. 8359–8388). https://proceedings.mlr.press/v235/chiang24b.html
Ding, D., Mallick, A., Wang, C., Sim, R., Mukherjee, S., Ruhle, V., Lakshmanan, L. V. S., & Awadallah, A. H. (2024). Hybrid LLM: Cost-efficient and quality-aware query routing [Preprint]. arXiv. https://arxiv.org/abs/2404.14618
Google. (n.d.). Rate limits. Retrieved July 17, 2026, from https://ai.google.dev/gemini-api/docs/rate-limits
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., & Steinhardt, J. (2021a). Measuring massive multitask language understanding. In International Conference on Learning Representations. https://openreview.net/forum?id=d7KBjmI3GmQ
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., & Steinhardt, J. (2021b). Measuring mathematical problem solving with the MATH dataset. In Advances in Neural Information Processing Systems (Vol. 34, pp. 3057–3074). https://proceedings.neurips.cc/paper/2021/hash/be83ab3ecd0db773eb2dc1b0a17836a1-Abstract.html
Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., Zhang, Y., & Ginsburg, B. (2024). RULER: What’s the real context size of your long-context language models? In First Conference on Language Modeling. https://arxiv.org/abs/2404.06654
Hu, Q. J., Bieker, J., Li, X., Jiang, N., Keigwin, B., Ranganath, G., Keutzer, K., & Upadhyay, S. K. (2024). RouterBench: A benchmark for multi-LLM routing system [Preprint]. arXiv. https://arxiv.org/abs/2403.12031
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., & Stoica, I. (2023). Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles. https://doi.org/10.1145/3600006.3613165
Li, C., Liu, G., & Zhao, Z. (2026). Cost-aware LLM-style routing for AIOps log analysis: Log parsing, anomaly detection, fault diagnosis, and incident summarization on LogEval task files. Journal of Technology Informatics and Engineering, 5(2), 91–103. https://doi.org/10.51903/jtie.v5i2.538
Li, H., Zhang, Y., Guo, Z., Wang, C., Tang, S., Zhang, Q., Chen, Y., Qi, B., Ye, P., Bai, L., Wang, Z., & Hu, S. (2026). LLMRouterBench: A unified benchmark for robust and efficient LLM routing [Preprint]. arXiv. https://arxiv.org/abs/2601.07206
Lu, K., Yuan, H., Lin, R., Lin, J., Yuan, Z., Zhou, C., & Zhou, J. (2024). Routing to the expert: Efficient reward-guided ensemble of large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 1964–1974). https://doi.org/10.18653/v1/2024.naacl-long.109
Mohammadshahi, A., Shaikh, A. R., & Yazdani, M. (2024). Routoo: Learning to route to large language models effectively [Preprint]. arXiv. https://arxiv.org/abs/2401.13979
Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., & Stoica, I. (2025). RouteLLM: Learning to route LLMs from preference data. In International Conference on Learning Representations. https://openreview.net/forum?id=8sSqNntaMr
OpenAI. (n.d.). Rate limits. Retrieved July 17, 2026, from https://developers.openai.com/api/docs/guides/rate-limits
Shah, S., & Shridhar, K. (2025). Select-then-route: A two-stage framework for efficient LLM routing. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track (pp. 425–441). https://doi.org/10.18653/v1/2025.emnlp-industry.28
Shen, Y., Liu, Y., Huang, Z., Yin, R., Zheng, X., & Huang, X. (2025). SATER: Cost-aware and latency-efficient LLM routing through speculative answer generation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 10515–10529). https://doi.org/10.18653/v1/2025.emnlp-main.531
Shnitzer, T., Ou, A., Silva, M., Soule, K., Sun, Y., Solomon, J., Thompson, N., & Yurochkin, M. (2023). Large language model routing with benchmark datasets [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2309.15789
Uddin, I., & Bauer, A. (2026). Conformal LLM routing with distribution-free safety guarantees. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop (pp. 791–799). https://doi.org/10.18653/v1/2026.acl-srw.70
Varangot-Reille, C., Bouvard, C., Gourru, A., Ciancone, M., Schaeffer, M., & Jacquenet, F. (2025). Doing more with less: A survey on routing strategies for resource optimisation in large language model-based systems [Preprint]. arXiv. https://arxiv.org/abs/2502.00409
Wang, X., Liu, Y., Cheng, W., Zhao, X., Chen, Z., Yu, W., Fu, Y., & Chen, H. (2025). MixLLM: Dynamic routing in mixed large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 10912–10922). https://doi.org/10.18653/v1/2025.naacl-long.545
Xin, Q. (2025). Hybrid cloud architecture for efficient and cost-effective large language model deployment. Journal of Information Systems and Informatics, 7(3), 2182–2195. https://doi.org/10.51519/journalisi.v7i3.1170
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena [Preprint]. arXiv. https://arxiv.org/abs/2306.05685
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., & Zhang, H. (2024). DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. In 18th USENIX Symposium on Operating Systems Design and Implementation (pp. 193–210). https://www.usenix.org/conference/osdi24/presentation/zhong-yinmin
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Lily Peng

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.