yi-rag

Retrieval for a historical corpus in Yiddish.

  • Historical research
  • RAG
  • Low-resource language
  • Cross-lingual retrieval

yi-rag is a research prototype I developed with historian Walter Koppmann to explore a Yiddish-language historical corpus from the Yiddish Book Center.

It explores how researchers can search across languages and find relevant passages while keeping the original sources available for inspection. Working with historical Yiddish texts raises questions about OCR quality, translation and what retrieval systems recognize—or miss—as relevant.

Translation can make a collection easier to search, but it can also change which passages surface and how historical language is represented. The prototype lets us explore these tradeoffs through different retrieval approaches, returning passages that researchers can examine in the original text.

A Yiddish-language expert provided insights into the limitations of machine translation from Yiddish. All machine-translated texts are accompanied by a disclaimer.

We presented the project as part of “AI and new technologies in the Global South: potentials and limits for history research” at the 2025 IALHI conference in Amsterdam.

GitHub repository · IALHI conference & programme