Web Codegen Scorer is a tool for evaluating the quality of web code generated by Large Language Models (LLMs).
You can use this tool to make evidence-based decisions relating to AI-generated code. For example:
- 🔄 Iterate on a system prompt to find most effective instructions for your project.
- ⚖️ Compare the code quality of code produced by different models.
- 📈 Monitor generated code quality over time as models and agents evolve.
Web Codegen Scorer is different from other code benchmarks in that it focuses specifically on web code and relies primarily on well-established measures of code quality.
- ⚙️ Configure your evaluations with different models, frameworks, and tools.
- ✍️ Specify system instructions and add MCP servers.
- 📋 Use built-in checks for build success, runtime errors, accessibility, security, LLM rating, and coding best practices. (More built-in checks coming soon!)
- 🔧 Automatically attempt to repair issues detected during code generating.
- 📊 View and compare results with an intuitive report view