
Project: Machine Learning for Improving Construction Carbon Estimations
Ponsu Muthuraman ’26 evaluated the use of a machine learning tool to improve construction carbon emission estimation.
Stanford’s Research Computing Facility (SRCF) operates two major computing clusters, the Sherlock HPC cluster and the Marlowe GPU cluster, which collectively support thousands of research jobs across campus. As AI and data-intensive research grow, so does the energy demand of this infrastructure. Despite Stanford’s broader sustainability commitments, the carbon footprint of research computing remains difficult to quantify and largely unmeasured.
Researchers had no visibility into the carbon cost of their jobs, no live estimates, or feedback, despite it being such a widely used resource. Looking at the future, SRCF manages infrastructure for reliability and performance, not emissions, and as Stanford planned major compute expansion (a colocation pilot, cloud, new facilities), it was unclear whether sustainability was being factored into those decisions at all.
In addition, no unified emissions accounting framework existed for Stanford’s research computing infrastructure. Energy use, hardware lifecycle data, and grid carbon intensity had never been consolidated. Translating raw power draw into carbon meant navigating Scope 2 (operational electricity) versus Scope 3 (embodied hardware carbon), alongside average versus marginal emission factors and how to credit renewables.
This fellowship project was motivated by the need to build a rigorous, reproducible framework for understanding and eventually reducing those emissions, connecting the dots between server racks, electricity consumption, and Stanford’s climate goals.
Explored smarter resource use on the Sherlock computing cluster to make these impacts transparent to users
The project began with a literature and standards review (GHG Protocol, LBNL data-center research, and HPC carbon accounting work from BU, UChicago, and Google) to ground the framework. Structured sacct exports and machine inventories were cleaned and analyzed to map existing data and gaps and to compute an energy and emissions baseline. The baseline emissions were derived by pairing per-job energy with historical grid carbon intensity (ElectricityMaps) and marginal signals (WattTime). Energy estimation was done through per-user and hardware-class models benchmarked against measured energy. Stakeholder engagement spanned Brian Chivers, Brian Hudspeth, Professor Carlos Diaz-Marin , Ben Rogers, SLAC’s Mark Yoder, and this project’s advisory council, via discussions, support, and feedback.
The project produced a structured emissions framework and a baseline, over a representative seven-month window of the jobs run on Sherlock, ~107,000 jobs consumed ~821 MWh and emitted ~152 t CO2e at a weighted-average grid intensity of 185 gCO2e/kWh. Deliverables include a live dashboard (baseline plus pre-run energy predictions, surfacing the facility PUE of 1.287), a FastAPI energy-estimation service, a per-user shell script researchers can run at submission time, and an automated per-job energy “receipt” added to existing PI usage emails. A researcher survey measured awareness. Conversations were started around shaping future expansion of compute, including tradeoffs between potential cloud providers and requiring providers to report job-level energy and carbon emissions. Key data gaps were identified, including job-level energy attribution and the marginal carbon intensity at SLAC.
The clearest lesson was that emissions accounting in an operational computing environment is as much an organizational challenge as a technical one. Progress depended on building trust with infrastructure staff and researchers, and on framing sustainability in terms that fit their existing priorities of reliability, cost, and performance. Very little of this work exists at the university level, so a rigorous, reproducible model at Stanford can serve as a template for other institutions, an outsized opportunity given how large research computing is now across academia. Professionally, the fellowship deepened the student fellow’s technical fluency in data-center energy systems and experience navigating institutional sustainability work, working with and learning from many players, IT operators to PIs relying on these systems daily.
Closing data gaps via stronger SRCF and SLAC partnerships. The framework can feed a public “Research Computing” category in Sustainable Stanford’s Data Hub. The dashboard, shell script, and email receipt can be iterated upon according to feedback. Shaping cloud compute procurement so vendors disclose job-level energy and carbon, and once compute capacity is expanded, carbon-aware scheduling simulations would be interesting to look into (California’s solar-driven grid swings make this especially promising).
PRIMARY PARTNER: SLAC Accelerator National Laboratory
Danica is studying CS (B.S. ’28). She has spoken at COP28 and been featured by the Chicago Tribune, New York Times, and NPR. She is passionate about building technologies that win on their own merits, faster, more innovative, more efficient, scaling fast and cutting emissions as a result. Last summer she was the sole hire, doing engineering at an a16z-backed energy and industrial-development startup. On campus, she co-directs Stanford Climate Week and the Stanford Ventures Frontier Fellowship.
Professor Carlos Diaz-Marin is leading a research group on data center heat waste water reuse.

Ponsu Muthuraman ’26 evaluated the use of a machine learning tool to improve construction carbon emission estimation.

Students are transforming reusable mugs into a community ritual–reducing waste while rebuilding campus culture.

Rachel Porter, PhD candidate, investigated alternatives for reducing reliance on Stanford’s natural gas-fired steam plant.