Număr curent
Revista Română de Informatică și Automatică / Vol. 36, Nr. 3, 2026
Operational efficiency in Large Language Models: A token optimization approach for high-volume data/query systems
Adrian BESIMI, Xhemal ZENUNI, Florije ISMAILI, Nuhi BESIMI, Artir DAUDI
This paper presents a two-layer token optimization approach applied in AI@SEEU: AI for Staff. The approach focuses on both the LLM Tool Selector input and the output structure of MCP tools. The aim of this work is to reduce token consumption in a private institutional deployment while preserving the information required to answer academic staff queries. The first optimization targets the Tool Selector prompt, which initially included extensive tool descriptions and redundant schema definitions that increased input token consumption. The second optimization restructures the output of internal academic tools by removing repeated contextual metadata and introducing a compact hierarchical format. The approach was evaluated across eight representative academic queries and resulted in an overall token reduction of 54.30%. The results demonstrate substantial reductions in token consumption; however, functional correctness, latency, cost, and generalizability are discussed as areas that require further validation.
Cuvinte cheie:
Large Language Models, Token optimization, Tool selection, MCP tools, AI.
Vizualizează articolul complet:
ACEST ARTICOL SE CITEAZĂ ASTFEL:
Adrian BESIMI,
Xhemal ZENUNI,
Florije ISMAILI,
Nuhi BESIMI,
Artir DAUDI,
„Operational efficiency in Large Language Models: A token optimization approach for high-volume data/query systems”,
Revista Română de Informatică și Automatică,
ISSN 1220-1758,
vol. 36(3),
pp. 67-81,
2026.
https://doi.org/10.33436/v36i3y202605