openai / evals

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Other
14.35k stars 2.54k forks source link

Add Gemini Solver #1503

Closed ojaffe closed 3 months ago

ojaffe commented 3 months ago

Adds a solver for Gemini 1.5 Pro. Stacked on #1501 and #1482. Using the solver requires the GEMINI_API_KEY environment variable

Test with:

oaieval generation/direct/gemini-pro bugged_tools