
Luís F. Gomes, Xin Zhou, Vincent Hellendoorn, Jonathan Aldrich, Rui Abreu, David Lo
Preprint — under review 2026
JupyterDraw automatically generates informal diagrams from Jupyter notebooks to help reviewers catch silent flaws (data leakage, improper evaluation) in AI-generated code. A controlled experiment with 26 participants found the visual overview caught +0.50 more flaws per task than a textual one, raising the share of reviews catching both planted flaws from 29% to 67%.
Luís F. Gomes, Xin Zhou, Vincent Hellendoorn, Jonathan Aldrich, Rui Abreu, David Lo
Preprint — under review 2026
JupyterDraw automatically generates informal diagrams from Jupyter notebooks to help reviewers catch silent flaws (data leakage, improper evaluation) in AI-generated code. A controlled experiment with 26 participants found the visual overview caught +0.50 more flaws per task than a textual one, raising the share of reviews catching both planted flaws from 29% to 67%.

Luís F. Gomes, Xin Zhou, David Lo, Rui Abreu
International Conference on AI Foundation Models and Software Engineering (FORGE) 2026
A vision paper introducing Visual Loop, a continuous visual development environment that keeps code and informal sketches in bidirectional synchronization. The prototype connects a code editor with a tablet-based visualization workspace, where freehand sketches are interpreted by multimodal LLMs grounded in static analysis and visual context — turning sketching from passive documentation into an active interface for software evolution.
Luís F. Gomes, Xin Zhou, David Lo, Rui Abreu
International Conference on AI Foundation Models and Software Engineering (FORGE) 2026
A vision paper introducing Visual Loop, a continuous visual development environment that keeps code and informal sketches in bidirectional synchronization. The prototype connects a code editor with a tablet-based visualization workspace, where freehand sketches are interpreted by multimodal LLMs grounded in static analysis and visual context — turning sketching from passive documentation into an active interface for software evolution.

Xin Zhou, Kisub Kim, Ting Zhang, Martin Weyssow, Luís F. Gomes, Guang Yang, Kui Liu, Xin Xia, David Lo
CORE A* International Conference on Automated Software Engineering (ASE) 2025
SE-Jury is an LLM-as-ensemble-judge metric that combines five distinct evaluation strategies to assess the functional correctness of generated software artifacts, with a dynamic team selection mechanism that picks the optimal subset of evaluators — cutting LLM API costs by ~50% while maintaining evaluation accuracy.
Xin Zhou, Kisub Kim, Ting Zhang, Martin Weyssow, Luís F. Gomes, Guang Yang, Kui Liu, Xin Xia, David Lo
CORE A* International Conference on Automated Software Engineering (ASE) 2025
SE-Jury is an LLM-as-ensemble-judge metric that combines five distinct evaluation strategies to assess the functional correctness of generated software artifacts, with a dynamic team selection mechanism that picks the optimal subset of evaluators — cutting LLM API costs by ~50% while maintaining evaluation accuracy.

Luís F. Gomes, Xin Zhou, David Lo, Rui Abreu
arXiv preprint 2025
An agentic LLM system that generates high-level visual documentation from source code, paired with AutoSketchEval — a reference-free evaluation framework (inspired by autoencoder reconstruction) that scores diagram quality with no ground truth, reaching AUC > 0.87 across 1,000 Jupyter notebooks.
Luís F. Gomes, Xin Zhou, David Lo, Rui Abreu
arXiv preprint 2025
An agentic LLM system that generates high-level visual documentation from source code, paired with AutoSketchEval — a reference-free evaluation framework (inspired by autoencoder reconstruction) that scores diagram quality with no ground truth, reaching AUC > 0.87 across 1,000 Jupyter notebooks.

Luís F. Gomes, Jonathan Aldrich, Rui Abreu, Vincent Hellendoorn
CORE A* International Conference on Software Engineering (ICSE) 2025
A VSCode assistant that turns hand-drawn ML workflow sketches into runnable Jupyter notebooks (79% structural accuracy, 49% reduction in coding), with a 19-participant developer study and an automated LLM-as-judge pipeline benchmarking sketch-to-code across GPT-4o, Gemini Pro, and Claude.
Luís F. Gomes, Jonathan Aldrich, Rui Abreu, Vincent Hellendoorn
CORE A* International Conference on Software Engineering (ICSE) 2025
A VSCode assistant that turns hand-drawn ML workflow sketches into runnable Jupyter notebooks (79% structural accuracy, 49% reduction in coding), with a 19-participant developer study and an automated LLM-as-judge pipeline benchmarking sketch-to-code across GPT-4o, Gemini Pro, and Claude.

Luís F. Gomes
SPLASH — Doctoral Symposium 2023
Doctoral symposium paper outlining a research agenda for visual sketching as an interface for machine learning development.
Luís F. Gomes
SPLASH — Doctoral Symposium 2023
Doctoral symposium paper outlining a research agenda for visual sketching as an interface for machine learning development.

Luís F. Gomes, César Analide, Elisabete Freitas
International Conference on Distributed Computing and Artificial Intelligence (DCAI) 2021
Neural network models for detecting distress and defects in road pavement imagery.
Luís F. Gomes, César Analide, Elisabete Freitas
International Conference on Distributed Computing and Artificial Intelligence (DCAI) 2021
Neural network models for detecting distress and defects in road pavement imagery.