WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
Finding in the paper: The benchmark provides 647 interaction test cases with explicit operations and expected outcomes. Its best reported tested combination reached 27.8% accuracy on those cases.
Limit: This result is benchmark- and system-specific; it is not an estimate of the success rate of every model or real-world application.
WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
Finding in the paper: The paper evaluates screenshot feedback, GUI-agent checks, backtracking, and select-best mechanisms; its ablation reports a measurable functional-testing contribution from GUI-agent feedback.
Limit: The reported results use the authors' benchmark and pipeline and do not guarantee that a separate product's evaluator is reliable.
From Runnable to Shippable: Multi-Agent Test-Driven Development for Generating Full-Stack Web Applications from Requirements
Finding in the paper: In a controlled study across 20 web applications, multiple backbones, two coding agents, and three TDD implementations, the authors report 15.5 to 23.7 percentage-point accuracy gains for sufficiently capable backbones. They report protocol tradeoffs by model capability, tester reliability, and deployment objective; Whole-Project TDD was most cost-efficient in most configurations.
Limit: These results describe the authors' tested applications, models, agents, and protocols. They do not establish that other evaluators or repair loops will be effective or cost-efficient.
PreMaQ: Predicting Maintainability-Related Quality of LLM-Generated Code Before Generation
Finding in the paper: In 2,695 Python benchmark tasks spanning four models, the authors report mean Spearman correlations of 0.57 for code-smell score and 0.65 for maintainability index. They use these signals to rank model outputs that are already functionally correct.
Limit: The study concerns Python tasks and specific open-weight models. Predicting maintainability does not replace running a generated web application or checking its user task.
Practical workflow
- Start from a specific audience task supported by research. Treat public issue and discussion records as leads, not proof of a broad need.
- Ask the model for the smallest typed AppSPEC product contract: purpose, user task, evidence references, and field definitions.
- Validate the contract, then compile it through the existing AppSPEC primitive catalog. Keep page chrome, arbitrary markup, scripts, and CSS outside model authority.
- Open the compiled module in the existing sandbox runner and use browser events to check the expected state change, persistence, and error-free behavior.
- Inspect the rendered result at desktop and mobile sizes, including alignment with the canonical shell and meaningful empty, success, and error states.
- Measure privacy-conscious adoption and task outcomes with explicit denominators. Use feedback as a new research lead and retain only what the published privacy policy permits.
Scope
These are summaries of the linked studies, not peer review. Their findings are specific to the reported methods and do not establish demand for Programmati.ca or predict product adoption.