Giovanni D'Antonio
Giovanni D'Antonio avatar image

Giovanni D'Antonio

AB. in Stats & CS and M.Sc. in CS @ Harvard || Tedx Speaker || Myllenium Award (2022) || Nova 111 List U25 in CS (2023)

Technical Projects Spotlight


SENNS AND ADVERSARIAL ROBUSTNESS
Valerio Pepe, Giovanni D’Antonio, Martin Dimitrov
School of Engineering and Applied Sciences
Harvard University
Cambridge, MA 02138
{valeriopepe, giovannidantonio, martin_dimitrov}@college.harvard.edu
December 10, 2024
ABSTRACT
Self-Explaining Neural Networks (SENNs, [ 1 ]) promise interpretable, intrinsically explainable
modeling by decomposing predictive processes into understandable concepts and relevance weights.
In this paper, we present a systematic investigation into the adversarial robustness of SENNs under
the Fast Gradient Sign Method (FGSM, [2 ]). We find that while SENNs can withstand minor,
unstructured noise with minimal performance degradation, more targeted “chunked” perturbations
severely compromise the model’s accuracy. Our analyses reveal distinct vulnerability profiles: the
aggregator exhibits non-monotonic accuracy patterns explained by its manipulation of dot products.
In contrast, the conceptizer and parameterizer display more predictable yet ultimately devastating
accuracy declines due to the aforementioned vulnerabilities to concept-based noise, which we term
the viable alternative hypothesis. We conclude by outlining opportunities for future work, including
exploring stronger attacks and more robust defenses, and checking if these results hold for non-
gradient-based adversarial attacks. Code for the paper can be found in the ‘report.ipynb’ file at
https://anonymous.4open.science/r/2822-SENN-Project-9D0E1
1 Introduction
Over the last decade, modern machine learning techniques have vastly improved on previous state-of-the-art compu-
tational performance across a variety of different tasks and modalities (from image recognition to text generation).
The main edge these techniques have over previous ones – their exploitation of complex nonlinear relationships
between representations – allows for superior modeling capabilities and greater flexibility in learning, both of which are
important desiderata for any successful machine learning model. However, this additional performance is not without
its drawbacks: deep nonlinear relationships between neurons also make the interpretation of the model’s inner workings
much harder, spurring the birth of machine learning interpretability as an academic discipline.
Most techniques for ML interpretability are post-hoc, focusing on inferring the best possible explanation for a certain
output given some feature of the model (its gradient, its performance on similar perturbed inputs, etc.). These types of
explanations are easy to produce and do not compromise on the model’s expressivity, but have one fatal flaw; if they are
meant for human consumption, explanations can be unfaithful to the actual workings of the model but still convince
human operators, potentially leading to the wrong outcome being applied to an input (e.g. the wrong diagnosis in the
medical field, the wrong sentence in legal applications, a loan denial in financial ones).
On the other hand, models that use explanations as part of their decision-making process do not have such risks: a
manipulated explanation would lead to the wrong outcome for the model, so unfaithful explanations are harder to
change to the benefit of a third party. However, issues can still arise if these models are vulnerable to adversarial
attacks on the input: if an input can be changed in a manner imperceptible to humans and lead to a different explanation
1This is a fork of the base repository we took the SENN implementation from, hence why the README has other people’s
names on it. We also use anonymous.4open.science to avoid having to un-private the GitHub Repo, since that is usually considered a
violation of the Honor Code.
Hierarchical Distributed Low-Communication
(HeDiLoCo) Training
Final Report
Romeo Dean
Harvard John A. Paulson School of Engineering and
Applied Sciences
Cambridge, MA, USA
Giovanni M. D’Antonio
Harvard John A. Paulson School of Engineering and
Applied Sciences
Cambridge, MA, USA
ABSTRACT
Frontier AI models are growing to multiple trillions of param-
eters, pushing companies to invest billions into new datacen-
ter construction and accelerator purchases. As AI companies
strive to keep up with current scaling trends [6 ], they are ex-
pected to push well past the capacity of any single datacenter
[ 5 ]. This trend forces developers to consider multi-datacenter
training strategies. However, geographically distributed dat-
acenters introduce high-latency, low-bandwidth intercon-
nects that make traditional synchronous training—requiring
global parameter aggregation at every step—prohibitively
expensive.
We present Hierarchical Distributed Low-Communication
(HeDiLoCo) training, a framework inspired by DiLoCo [ 1 ],
that extends asynchronous training with flexibility for hier-
archical topologies. Workers in the same campus can syn-
chronize frequently at low latency, while synchronization
across distant regions can occur less frequently, to efficiently
balance communication overhead with model convergence
speed.
Our small experiments with a 100K parameter Transformer
on the tiny Shakespeare dataset demonstrate that HeDiLoCo
can speedup training times relative to a synchronous
baseline by 100 times while only incurring a 1-2% final
validation loss penalty. Our results highlight HeDiLoCo’s
promise as a flexible, scalable, cost-effective framework for
future multi-trillion parameter, multi-datacenter AI training.
1 INTRODUCTION
Frontier Artificial Intelligence (AI) models now contain tril-
lions of parameters, leading to training costs that can exceed
$1 billion. These exponential scaling trends are driven by
intense AI company competition and the rapid growth of
model capabilities. However, traditional compute colocation
strategies face bottlenecks, such as limited power availabil-
ity, construction timelines, and the inability to build new
datacenters fast enough to meet demand. For example, Mi-
crosoft recently announced a $7 billion investment in fiber
cabling to interconnect its AI-dedicated infrastructure across
regions [7].
A natural solution to these challenges is distributed train-
ing across geographically dispersed data centers. Unfortu-
nately, this approach introduces formidable issues: high-
latency and bandwidth-limited interconnects make tradi-
tional synchronous training—requiring global parameter ag-
gregation at every step—prohibitively inefficient. To address
these limitations, we propose Hierarchical Distributed
Low-Communication (HeDiLoCo) training, a novel frame-
work that reduces communication overhead by leveraging
hierarchical synchronization structures tailored to real-world
data center topologies.
We propose Hierarchical Distributed Low Communi-
cation (HeDiLoCo) training, an approach that extends the
asynchronous concept introduced by DiLoCo [ 1 ] with a hier-
archical structure. While DiLoCo reduces communication by
synchronizing every 𝐻 steps, HeDiLoCo further differenti-
ates between intra-campus (lower latency) and inter-campus
(higher latency) connections. Workers within the same re-
gion synchronize frequently, while workers across distant re-
gions synchronize less frequently. This hierarchical approach
can greatly reduce communication costs while maintaining
robust model convergence, reflecting real-world landscapes
where datacenters cluster into campuses and states.
1.1 Problem Definition
The exponential scaling trends in Artificial Intelligence (AI)
models, driven by fierce competition among AI companies,
have resulted in training costs surpassing $1 billion and
compute demands that exceed the capacity of any single
datacenter. Traditional colocation strategies face significant
bottlenecks, including power limitations, construction de-
lays, and the logistical challenges of building new facilities
quickly enough to meet demand. Distributed training offers
a natural solution but introduces new challenges.
The key challenges we address include:
(1) Minimizing communication overhead: High-latency,
bandwidth-limited links between geographically dis-
tributed data centers make global synchronization
1
Hierarchical Distributed Low-Communication (HeDiLoCo) Training
Incorporating Unspent Funds in Participatory Budgeting
A Statistical Analysis on Food Deserts (Top Team)

Non-Academic Experience


A quick Intro about me

My name is Giovanni and I am a rising Senior at Harvard studying an A.B. in Statistics + Computer Science and a concurrent S.M. in Computer Science. I am from a small town in southern Italy and graduated from a public high school in the province of Naples, where the high school dropout rate is more than twice the European average.

From a young age, I had to self-learn English and Mathematics, with 70% of my teachers not even having a Bachelor's degree. Then, I fell in love with Computer Science during my Freshman Year at Harvard and went all the way to graduate courses and research. I am also one of the only 10 Harvard College Students selected as a Robert family fellow for potential innovators, being able to study MBA classes. I am a 2 times gold medal in philosophy olympiad and have worked with the Italian parliament in the past

I love game theory, hiking, AI safety, and philosophy.

Presidential Speech | National TV
TEDx | Probabilistic Thinking
Lecture | University of Salerno

Resume & Work Experience


Resume | Giovanni D'Antonio
January 2024 - Quant Research
January 2024 - Quant Research
Spring 2023 - Consultant Product Team
Spring 2023 - Consultant Product Team

📋 Honors & Awards

  • Recognized as one of the most disruptive 20 Italians U30 in 2022 at MylleniumAwards

  • Ranked among top 10 Italian Computer Science talents Under 25 (By Nova Talent)

  • Winner of Harvard Center of European Studies Grant Award Summer 2023

  • Represented Italian Students at "Procida Capitale della Cultura" presidential ceremony

  • Received Rotary Club's Paul Harris Medal for social service at 18 (Youngest in history)

  • Gold Medalist in Philosophy Olympiads and 4th at the IPO in 2021 & 2022

  • Chosen as one of the 61 high school students in Italy for "I Fuoriclasse della Scuola" aimed to prize the merit of the highest achieving students in Italy by Olympiads medal (First in History to Win this Award for two consecutive years)

🛠 Skills & Hobbies

  • Programming: Python, R, C, SQL, Java, JavaScript, HTML, CSS, OCaml, MATLAB, TensorFlow, PyTorch, RedShift (AWS), Excel

  • Languages: Italian (native), English(advanced/bilingual), Spanish (beginner)

  • Interests: Hiking, Backpacking, Linguistics, Philosophy, Poker

Technical Projects


Coursework


Transcript
philosophy 1
A Spatially-Aware Search Engine for Textual Content in Images
philosophy 2
CV___Giovanni_M__D_Antonio (4).pdf