Samenvatting
With the advent of foundation models and generative AI, especially the recent explosion in Large Language Models (LLMs), we see a whole new type of AI-enabled systems: LLM-based systems or LLM systems in short. Inspired by ChatGPT and its possibilities, many developers want to build there own chatbots, trained on their own set of documents, e.g. as an intelligent search engine. For this specific text generation task they have to 1) select the most appropriate LLM, sometimes fine-tune it, 2) engineer the document retrieval step (Retrieval Augmented Generation, RAG), 3) engineer the prompt, 4) engineer a user interface that hides the complexity of prompts and answers to end users.
Especially prompt engineering is a new activity introduced in LLM systems. It is intrinsically hard as the possibilities are endless, prompts are hard to test or compare, the result might vary with different LLM models or model versions, prompts are difficult to debug, you need domain expertise (and language skills!) to engineer fitting prompts for the task at hand, and so on. For LLMs however, prompt engineering is the main way the models can be adapted to support specific tasks.
So, where in previous work we concluded that AI-enabled systems are data + model + code, for LLM systems we must conclude that they are data + model + prompt + code. Where it must also be noted that with LLM systems the model is usually provided by an external party and thus hard or inpossible for the developer to control, other than by engineering prompts. The external party might however frequently update its LLM and this might necessitate a system update for the LLM system as well.
In our work, we analyze the quality characteristics of LLM systems and discuss the challenges for engineering LLM systems. We also present the solutions we have found untill now to address the quality characteristics and the challenges. In future work we will engineer more LLM systems together with our students and workfield partners, thereby adding to the body of knowledge on LLM engineering.
Especially prompt engineering is a new activity introduced in LLM systems. It is intrinsically hard as the possibilities are endless, prompts are hard to test or compare, the result might vary with different LLM models or model versions, prompts are difficult to debug, you need domain expertise (and language skills!) to engineer fitting prompts for the task at hand, and so on. For LLMs however, prompt engineering is the main way the models can be adapted to support specific tasks.
So, where in previous work we concluded that AI-enabled systems are data + model + code, for LLM systems we must conclude that they are data + model + prompt + code. Where it must also be noted that with LLM systems the model is usually provided by an external party and thus hard or inpossible for the developer to control, other than by engineering prompts. The external party might however frequently update its LLM and this might necessitate a system update for the LLM system as well.
In our work, we analyze the quality characteristics of LLM systems and discuss the challenges for engineering LLM systems. We also present the solutions we have found untill now to address the quality characteristics and the challenges. In future work we will engineer more LLM systems together with our students and workfield partners, thereby adding to the body of knowledge on LLM engineering.
| Originele taal | Engels |
|---|---|
| Status | Gepubliceerd - apr 2025 |
| Evenement | ICT.OPEN - Duur: 16 apr 2025 → 17 apr 2025 |
Congres
| Congres | ICT.OPEN |
|---|---|
| Periode | 16/04/25 → 17/04/25 |
Vingerafdruk
Duik in de onderzoeksthema's van 'LLMOps: Engineering trustworthy LLM Systems'. Samen vormen ze een unieke vingerafdruk.Citeer dit
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver