smallevals — CPU-fast, GPU-blazing fast offline retrieval evaluation for RAG systems with tiny QA models.
-
Updated
Dec 4, 2025 - Python
smallevals — CPU-fast, GPU-blazing fast offline retrieval evaluation for RAG systems with tiny QA models.
Local decision recorder and offline policy evaluation for coding agents. Advisory tooling with explicit evidence and authority boundaries.
Lightweight Python library for interactive demo and inspection of recommender systems in Streamlit.
Support-ticket dedup & resolution finder — pgvector semantic retrieval evaluated with Netflix XP-style offline replay (recall@k / MRR). Shares a domain-agnostic similarity engine with its sibling repo 'clause'. Next.js 15 · Supabase · HF embeddings.
Offline prototype for intent-aware queue adaptation in music recommendation systems
Recommend Signal — temporal offline evaluation for recommendation policies, with explicit causal boundaries.
CTR/ranking fundamentals practice with feature crossing, Logistic Regression baselines, AUC/LogLoss/nDCG notes and reproducible evaluation scripts.
Algorithm Intern Candidate | Recommendation / Search Retrieval / Ranking | PyTorch + Faiss | Offline Evaluation / Negative Sampling / Badcase Analysis
Offline RAG retrieval-quality harness. Recall@k, nDCG, MRR, chunking diagnostics, regression diffs. No LLM-as-judge required. CI-friendly.
可复现的中文离线内容推荐应用原型,覆盖合成行为数据、动态兴趣画像、双路召回、个性化排序、多样性重排、推荐解释与离线评估。
轻量级终端 AI Agent Harness:手写状态机、能力审批、Checkpoint 恢复与增量审计;仅 1 个直接运行依赖,500+ 项测试与 10 项离线 Evals。
Neural Thompson Sampling contextual bandit for personalized Type 2 Diabetes therapy selection — training pipeline, offline policy evaluation (IPS/SNIPS/DM/DR), safety gates, drift monitoring, and LLM-generated clinical explanations.
MovieLens 1M 推荐系统:无泄漏时间切分与 Top-K 评测 / Leakage-safe temporal splitting and Top-K recommender evaluation.
Auditable multimodal sequential recommendation research with exact full-catalog evaluation, frozen model-selection gates, uncertainty-aware slicing, routing diagnostics, robustness analysis, and reproducible evidence.
End-to-end joke recommendation system with offline evaluation, FastAPI serving, Docker, and model artifacts.
Budgeted high-fidelity evaluation for symbolic regression
Production-ready recommender system suite: serving API, pipelines, algorithm SDK, and evaluation tooling.
Controlled study of offline recommender model-selection stability across logging regimes and behavioral targets.
Spotify-style music discovery platform with Spotify OAuth, hybrid recommendation, ALS, Word2Vec-style embeddings, explainable recommendations, and offline evaluation.
Experimental study of pCTR prediction, calibration, value-aware ranking, and auction decision robustness using real RTB logs.
To associate your repository with the offline-evaluation topic, visit your repo's landing page and select "manage topics."