How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
Papers with Code was revived three months ago to make open AI research accessible and digestible, helping people find paper artifacts, track state-of-the-art results, and build on each other's work. Its search engine must handle exact titles and arXiv IDs, but also fuzzy queries like “small language models for code generation,” navigational requests such as “the original BERT paper,” and incomplete or typo-ridden titles — while staying fast even when model services are cold or unavailable.
The system is built as a hybrid search engine, drawing on the team's prior RAG experience at ML6, because hybrid search typically outperforms keyword- or vector-only approaches. PostgreSQL full-text search provides a fast lexical baseline, pgvector adds dense embedding recall, and the reciprocal rank fusion (RRF) algorithm combines both result sets. Rerankers (cross-encoders) could improve results further, but they add overhead and latency.
Three Hugging Face services power the dense embeddings: HF Jobs provides burstable GPU compute for embedding the corpus, HF Storage Buckets offers durable handoff between the database, experiments, and Jobs, and HF Inference Endpoints serves low-latency embeddings for live queries and incremental updates. The system currently maintains embeddings for more than 110,000 papers from arXiv and Daily Papers, with the architecture deliberately split into an offline corpus build and an online search service.
The post walks through the architecture, design decisions, and production lessons, showing how the dual retrieval strategy gives both exact and semantic matches — letting humans and agents search via the website or the pwc CLI command exposed as a Skill.