LLM · RAG · Agentic Systems · Final Evaluation

Evidentia

A local scientific reading assistant that answers and compares arXiv papers using only retrieved, cited and inspectable evidence.

LangGraphQdrantFastAPIOpenVINODocker

Use case: literature review, comparison and evidence checkingExecution: local, modest CPU/GPUYear: 2026

Final v6 evaluation complete: 30 benchmark cases, all abstentions handled correctly
29/30End-to-end answers passed, or 96.7%.
100%Expected-document retrieval recall.
100%Citation precision and recall.
86.1%Mean factual coverage across answerable cases.
56.5 sMean latency per question on local CPU.

1. Evidence-first scientific RAG

Scientific assistants can produce a convincing answer without making its support easy to verify. Evidentia takes the opposite approach: every answer is built from evidence present in the selected corpus, linked to citations, and backed by the exact chunks supplied to the model.

Product question. How can a student, researcher or engineer interrogate several publications without losing factual attribution, especially when comparing methods, training objectives or architectural components?

The current prototype focuses on a curated set of computer-vision papers. Its interface lets users select sources, ask in English or French, and inspect retrieval, generation, citations and any abstention decision.

2. Local agentic pipeline

  1. IngestionarXiv PDFs are extracted with Docling and split into usable evidence passages.
  2. RetrievalMultilingual E5 embeddings in Qdrant, BM25 and reciprocal-rank fusion retrieve candidate evidence.
  3. OrchestrationLangGraph analyses the question, creates intent-specific facets and balances sources in comparisons.
  4. Controlled answerQwen 2.5 7B, quantized with OpenVINO, writes locally while a contract checks language, citations and facts.

Hybrid, multi-facet retrieval

For a complex question, the system combines dense and lexical retrieval, fuses rankings, then reranks evidence according to the detected intent: supervision, architecture, data, frozen components or comparison.

Fact planning and response contract

After evidence selection, a deterministic fact plan identifies what must be covered and which sources may support it. A response contract checks language, citation validity and factual completeness, with at most one corrective pass.

3. Usage examples

Understand a method

BLIP-2’s Q-Former

“Why is BLIP-2’s Q-Former described as an information bottleneck between vision and the LLM?”

Evidentia retrieves the relevant BLIP-2 passages, then explains how learnable queries extract useful visual information from a frozen image encoder before passing it to a frozen language model. The cited passages expose the architecture behind the answer.

Compare papers

CLIP and DINOv2

“Compare the training objectives of CLIP and DINOv2.”

The graph creates a balanced retrieval path per source, preventing CLIP’s contrastive objective from being attributed to DINOv2. The final answer separates CLIP’s image-text pairs from DINOv2’s teacher-student self-supervision.

What the interface makes inspectable

  • the selected papers and the chunks actually passed to the model;
  • citations with their source page and section;
  • the LangGraph trace: retrieval strategy, intent, fact plan, correction and contract status;
  • an explicit abstention when the available evidence is insufficient.

4. From failure modes to guardrails

The final benchmark’s remaining hard case asked how the LVD-142M dataset is curated without text metadata. The relevant paper was retrieved, yet the answer covered only part of the required mechanism. This distinction matters: finding the right paper is not enough if the decisive passage or detail is not carried into the answer.

Select the decisive evidence

A relevant citation does not guarantee that the selected passage contains the requested mechanism, number or component. Query facets, dense and lexical fusion, and intent-guided reranking focus the context on evidence that answers every part of the question.

Prefer transparency to a plausible answer

The fact plan links each expected element to its authorized sources. The contract checks language, citations and completeness; when a blocking constraint fails, the system abstains explicitly instead of presenting an unsupported answer as certain.

Separating retrieval, writing and validation makes system behavior inspectable. Users can tell apart missing evidence, an incomplete answer and a refusal triggered by the response contract.

Evaluation scope. Factual coverage is based on explicit term groups. This makes campaigns reproducible, but a human review remains necessary to judge scientific fidelity in nuanced answers.

5. Next steps

The next objective is to broaden the available evidence while preserving the same traceability: scientific figures and tables, followed by calibrated human evaluation on the most demanding answers.

  • Extend to scientific figures with a VLM that can connect text, tables, charts and images.
  • Add structured human review to complement deterministic attribution and factual-coverage metrics.
  • Explore SFT and preference alignment once evaluation criteria and retrieval are stable.