Vibe Coding Discover

Use Cases

315 use cases for “Eval”

Long-term Memory And Preference Retrieval

1 project

Long-term Memory With Vector + Keyword Retrieval

1 project

Manage Evaluators And Experiments For Agent Quality

1 project

Manage Skill Installs And Eval Coverage In A Desktop App

1 project

Mcp Server For Live Data Retrieval In Claude Code

1 project

Metadata-filtered Document Retrieval

1 project

Metadata-filtered Retrieval For Llm Context

1 project

Model Evaluation With Top-1/top-3 Accuracy And Expected Calibration Error

1 project

Monitor And Evaluate Agent Workflows

1 project

Monitor Token Usage, Latency, And Retrieval Spans

1 project

Multi-hop Retrieval And Question Answering

1 project

Multi-hop Retrieval And Recursive Query Decomposition Over Markdown Notes

1 project

Multi-source Document Retrieval

1 project

Multi-tenant Retrieval Where Callers Only Read Their Own Chunks

1 project

Naive Rag Mode For Plain Vector Retrieval

1 project

Offline Semantic Retrieval Over Canon Excerpts

1 project

Online Evaluation Rules For Live Traffic

1 project

Orchestrate Multi-agent Conversations With Retrieval

1 project

Pattern (geju) And Shensha Evaluation

1 project

Pk/pd Modelling And Dose-regimen Evaluation

1 project

Practice Evaluation And Reliability Engineering For Llms

1 project

Prepare Training/eval Corpora For Language Models

1 project

Production Agent Deployment With Tracing And Evaluation

1 project

Prompt Evaluation And Iteration

1 project

Prompt-injection-aware Tool Selection Evals

1 project

Prototype And Evaluate Agents With Langsmith

1 project

Provide Retrieval Data For Llm Pipelines

1 project

Query Engines With Retrieval And Reranking

1 project

Query Traces And Run Evals From Coding Agents Via Opik Mcp

1 project

R&d/ml Data Layer With Versioned Experiments For Training And Eval

1 project

Rag Evaluation With Ragas And Tracing With Langfuse

1 project

Rag Knowledge Base With Document Upload And Retrieval

1 project

Rag Knowledge Retrieval With Vector + Full-text + Scalar Filters

1 project

Rag Retrieval Backend For Llm Apps

1 project

Rag Retrieval Over Redis Vector Indexes

1 project

Rag-style Retrieval Over Accumulated Agent Knowledge

1 project

Reasoning-based Retrieval Over Long Documents Without A Vector DB

1 project

Red-team / Safety Evaluation Of Abliteration (advbench, Harmbench, Strongreject)

1 project

Reduce Token Usage Via Hybrid Memory Retrieval

1 project

Regression-test Coding Agents With A Git-worktree Eval Toolchain

1 project

Replace Vector-db Memory With Zero-dependency Local Lexical Retrieval

1 project

Reproduce Gaia Benchmark Agent Evaluations

1 project

Research Keywords And Evaluate Bids

1 project

Retrieval Over Large Documents With Configurable Embeddings

1 project

Retrieval-augmented Context For Llm Chats With Zero Api Calls

1 project

Retrieval-augmented Generation (rag) Pipelines

1 project

Retrieval-augmented Generation With Graph Context

1 project

Run A Zero-install 11-tool Stdlib Demo Mcp Server For Evaluation

1 project

Run Agent Eval Suites And Llm Observability Tracing

1 project

Run Agent Evaluation And Codex Parity Benchmarks

1 project

Run Ai-tool Routing Evals For Plugin Integration

1 project

Run And Evaluate Diffusion-model Sampling

1 project

Run Built-in Evals With Pass@k Metrics

1 project

Run Cross-client Evaluations Of Mcp Integrations

1 project

Run Evals Fully Locally To Keep Prompts Private

1 project

Run Evals That Score Agent Response Formatting

1 project

Run Evaluation Suites And Calibration Metrics For Scoring Pipelines

1 project

Run Large-scale Concurrent Agent Evaluations

1 project

Run Llm-as-a-judge And Code-based Evaluations

1 project

Run Llm-as-judge Response And Retrieval Evaluations

1 project

Run Local Eval Cases And Suites With Sqlite Telemetry

1 project

Run Reproducible Agent Evaluations As Experiment Cells And Attempts

1 project

Run Skill Evals And Test Harnesses In CI

1 project

Run Systematic Llm Evaluations With Built-in Judges And Metrics

1 project

Run Trigger And Output Eval Suites Against Generated Skills

1 project

Run Zero-cloud, No-api-key Memory Retrieval Locally Via Mcp Tools

1 project

Running Agent Evaluations

1 project

Running Bias, Toxicity And Fairness Evaluations

1 project

Running Local Vllm Vs Remote Api Evaluation Pipelines

1 project

Sandboxed Evaluation Of Untrusted Third-party Mcp Configurations

1 project

Score Rag Pipelines For Retrieval And Answer Quality

1 project

Screen Tool Calls Via Mcp Evaluation Servers

1 project

Script Vault Retrieval From The Cli With Filters And Json Output

1 project

Search A Local Knowledge Base With Semantic And Graph Retrieval

1 project

Search Knowledge With Semantic And Keyword Retrieval

1 project

Search Personal Material With Semantic And Text Retrieval

1 project

Search Project Knowledge With Sub-300µs In-memory Bm25 Retrieval

1 project

Search Repository Documents And Pdfs With Semantic Retrieval

1 project

Security Benchmark Evaluation (xbow/cybench)

1 project

Self-evaluate Rendered Output At Every Cut Boundary Before Preview

1 project

Self-host An Agent Trace And Eval Server With No User Code Executing On It

1 project

Self-query And Parent-document Retrieval

1 project

Semantic (lossless-compressed) Memory Retrieval To Cut Token Usage

1 project

Semantic + Keyword Retrieval Over Markdown Conversation History

1 project

Semantic Retrieval Of Past Context Via Vector Embeddings

1 project

Semantic Search With Hybrid Vector + Rerank Retrieval

1 project

Set Up Eval/verification Loops With Pass@k Metrics

1 project

Stock Price Retrieval And Disclosure Monitoring

1 project

Store And Query Document Embeddings For Rag Retrieval

1 project

Sub-millisecond Local Vector + Keyword Retrieval For Coding Agents

1 project

Swe-bench Verified Evaluation

1 project

Template-based Chunking And Retrieval Tuning

1 project

Testing Rag Pipelines For Pii Leakage And Cross-context Retrieval

1 project

Tool-gated Retrieval With Receipts For Auditing Agent Decisions

1 project

Trace And Debug Llm Calls, Retrieval And Agent Actions

1 project

Trace And Evaluate Agent Behavior

1 project

Trace And Evaluate Hosted Agents End-to-end

1 project

Trace And Evaluate Llm Apps With Langfuse

1 project

Trace And Evaluate Llm/agent Runs With Langfuse

1 project

Trace Retrieval And Vector-store Problems

1 project