InPars-Light: Cost-Effective Unsupervised Training of Efficient Rankers

Boytsov, Leonid; Patel, Pritash; Sourabh, Vivek; Nisar, Riddhi; Kundu, Sayani; Ramanathan, Ramya; Nyberg, Eric

doi:10.48550/arxiv.2301.02998

Cited by 1 publication

(1 citation statement)

References 33 publications

(57 reference statements)

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Data augmentation for information retrieval (IR) has gained attention as a promising research area. Previous studies use large language models (LLMs) (Zhao et al, 2023) to generate synthetic training data for retrievers (Jeronymo et al, 2023;Dai et al, 2023;Boytsov et al, 2023;Bonifacio et al, 2022), significantly improving effectiveness of unsupervised retrievers. Specifically, these studies all build pseudo query-document pairs by gen-erating synthetic queries given documents in an existing corpus.…”

Section: Introductionmentioning

confidence: 99%

Expand, Highlight, Generate: RL-driven Document Generation for Passage Reranking

Askari,

Aliannejadi,

Meng

et al. 2023

Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing

View full text Add to dashboard Cite

Generating synthetic training data based on large language models (LLMs) for ranking models has gained attention recently. Prior studies use LLMs to build pseudo querydocument pairs by generating synthetic queries from documents in a corpus. In this paper, we propose a new perspective of data augmentation: generating synthetic documents from queries. To achieve this, we propose DocGen, that consists of a three-step pipeline that utilizes the few-shot capabilities of LLMs. Doc-Gen pipeline performs synthetic document generation by (i) expanding, (ii) highlighting the original query, and then (iii) generating a synthetic document that is likely to be relevant to the query. To further improve the relevance between generated synthetic documents and their corresponding queries, we propose DocGen-RL, which regards the estimated relevance of the document as a reward and leverages reinforcement learning (RL) to optimize Doc-Gen pipeline. Extensive experiments demonstrate that DocGen and DocGen-RL significantly outperform existing state-of-the-art data augmentation methods, such as InPars, indicating that our new perspective of generating documents leverages the capacity of LLMs in generating synthetic data more effectively. We release the code, generated data, and model checkpoints to foster research in this area 1 .

show abstract