
Extended Reality (XR) is transforming the way industrial workers are trained and supported. New recruits can learn to operate a machine before working on the real one, new equipment can be explored in a virtual environment, and hazardous procedures can be rehearsed where a mistake carries no cost or risk.
Throughout these activities, workers need precise guidance, such as the steps of a procedure, the meaning of an alarm code or the limit that applies to a specific component. An AI assistant integrated into the XR environment can provide this guidance at the point of work. However, the required knowledge lies in hundreds of company manuals and technical procedures that are not publicly available, so general purpose language models cannot provide it reliably.
This is where retrieval augmented generation (RAG) plays a crucial role. The documentation is ingested into a searchable knowledge base, and every answer is generated from the passages retrieved from it, so that guidance remains faithful to the documentation. Within the XR5.0 project, Innov-Acts has developed a multimodal RAG assistant for industrial documentation and evaluated how each design decision affects its answer quality, cost and response time, as presented in our paper.
Industrial manuals combine technical text with diagrams, schematics, screenshots and parameter tables, and much of the information a worker needs is carried by these figures rather than by the text. A text-only assistant discards this material and leaves some questions unanswered. For this reason, Innov-Acts created a multimodal assistant. It processes figures together with text and returns them to the worker alongside the written answer.
During ingestion, each manual is converted into structured text, and every figure is passed to a vision language model, that creates a short description of the image. This description is appended to the document text at the position the figure occupied, which makes the figure searchable in the same way as the text. Each passage also keeps references to its figures, so the corresponding images can be returned whenever the passage is retrieved.
Each question is handled by three dedicated AI stages.
- Retrieval Planning. The planner receives the user question, chat history and a catalogue of the available documents in the database. It then decides in which documents it should search and what the query should be. Hybrid retrieval then combines dense and lexical retrieval over both the planner’s query and the worker’s original wording, so that exact terms such as alarm codes and part identifiers are always preserved.
- Figure Selection. A dedicated stage reviews the retrieved passages and their figures and selects those that should accompany the answer.
- Answer Generation. The writer produces an answer grounded in the retrieved text and the descriptions of the selected figures, and states when the documentation is insufficient.
The same pipeline serves unrelated industrial domains through configuration rather than redevelopment.
We evaluated six successive configurations, starting from a single function calling agent (FCA) baseline and ending with the deployed system. All configurations used the same index of 41 manuals from two unrelated industrial domains, pipe rehabilitation and bar feeder service, deliberately combined to increase its heterogeneity. The test set consists of 145 questions phrased as a technician would ask them, with handwritten ground truth answers. Answers were scored at claim level with the RAGChecker framework, and cost and response time were recorded for every configuration. Figure selection was evaluated separately on 22 questions against manually annotated figure sets. The box plots below show the scores of the baseline (V0) and the final deployed system (V5), while the full comparison of all six configurations is presented in the paper.

Key Results
- Improved Answer Quality. From the baseline to the deployed system, F1 increased from 60.7 to 76.0 and claim recall from 76.9 to 91.7.
- Stronger Grounding. Faithfulness to the retrieved documentation increased from 91.7 to 96.7, while hallucination decreased from 6.2 to 2.4.
- Robust Retrieval for Difficult Questions. The retrieval optimizations had the strongest effect on the weakest cases, with the lower end of the claim recall distribution rising from 17.7 to 80.0.
- Efficiency in Cost and Latency. The deployed system responds in 5.8 seconds compared with 7.0 seconds for the baseline, at approximately USD 2.40 per thousand queries. Separating planning from answer generation preserved answer quality while reducing cost and latency.
- Accurate Figure Selection. The assistant returned 92.1% of the relevant figures without any image specific retrieval optimization. The main remaining error is the inclusion of an additional figure.
- Reliable Abstention. All configurations correctly mentioned the lack of available information to the questions not covered by the documentation.
A key finding is that the largest single improvement came from the final instructions given to the planner and writer, which increased F1 by 7.5 points without additional latency while claim recall remained almost unchanged. Answer quality therefore depends not only on the evidence retrieved but also on how that evidence is used. Plotting answer quality against operating cost and response time shows that the deployed configuration lies on the Pareto frontier in both cases, combining the highest answer quality with a response time faster than the baseline. The same analysis provides a practical basis for selecting a configuration according to the budget and response time requirements of each deployment.

This work is closely aligned with the goals of the XR5.0 project, which aims to advance AI driven, human centric XR applications for Industry 5.0. The assistant can serve as the knowledge layer behind XR training and support, delivering grounded answers and the relevant figures at the point of work.
Our ongoing work within XR5.0 focuses on further improving figure selection and extending the assistant to additional pilot documentation. Innov-Acts thereby contributes to trustworthy AI support that helps industrial workers carry out their tasks more safely, accurately and efficiently.
