|
| ||
ResearchOak Ridge National Laboratory
Dates: Summer 2026
Mentor:
My work at Oak Ridge focused on finding a solution to a pesky and ongoing problem in retrieval-augmented generation (RAG) systems.
Say you're building a RAG system and have a stack of
documents you're using as your knowledge source. Now imagine a subset of those documents contains 'restricted' information, i.e., stuff you don't want
to send to OpenAI. Outside of those restricted documents, the rest of the dataset (perhaps the bulk of it) is safe to send.
Many teams in this scenario implement one of two solutions:
(1) Run all local inference.
(2) Rely on contractual guarantees, for example, an OpenAI promise to "not train" on your data.
(1) can be viable, but requires serious infrastructure investment to run large open-weight models. The current largest open-weight models
are over a trillion parameters. So, unless your team can afford potentially hundreds of thousands of dollars in infrastructure costs upfront to run them efficiently, you're left working with
smaller models. And of course, if data security is an absolute requirement for your team, (2) might be completely off the table.
Our group proposed a third solution: dynamically select between local and cloud inference depending on the security
classification of the retrieved evidence. If the evidence contains no restricted information, go ahead and send the answer generation request to the cloud and reap the benefits
of higher answer quality. Otherwise, send the request to local compute.
I built a full-stack application with a PostgreSQL vector database to demonstrate the utility of this idea, and tested it on a dataset of Department of
Energy project documents that were prelabelled as either 'restricted' or 'releasable'. The benefits of this system were most pronounced on datasets in which most of the documents were labelled 'releasable'. On these
datasets, our system kept all 'restricted' evidence local while improving average answer quality over all-local inference.
A comprehensive write-up of this approach is under review for the Privacy and Security of Big Data Special Session at IEEE Big Data 2026.
____________________________________________________________________
University of Michigan
Dates: January 2026 — Present
Faculty Mentor:
Industry Mentors:
This year-long program is a collaboration between University of Michigan student researchers and
Walbridge, a Detroit-based construction company specializing in
large-scale industrial and manufacturing projects. My team and I work directly with the technology leaders at Walbridge to build
an AI chat service using retrieval augmented generation (RAG) on their sensitive project documents.
I work on answer verification, data cleaning, and agent architecture. Most of our
solutions involve a combination of MS Copilot Studio, Power Automate workflows, and a data maintenance service, like SharePoint, Dataverse,
or Excel depending on the application.
____________________________________________________________________
University of Tennessee, Knoxville
Dates: Summer 2024
Mentors:
The objective of this project was to investigate the performance differences between heuristic programming methods and
RL methods in vehicle traffic simulations. We were specifically interested in designing simulations that involved a 'priority' vehicle, e.g.,
a police car or ambulance, as it maneuvered through an environment of autonomous vehicles. We imagined a vehicle-to-vehicle scenario where the priority
car could communicate its presence to the network of autonomous vehicles, prompting them to enter a 'safe-driving' mode to avoid
obstructing the priority car. This safe-driving mode featured a sequence of actions such as slowing down, pulling to the side, or clearing a path,
which were either heuristic or RL based.
All simulations were created using the open-source traffic simulation platform SUMO and
controlled via the TraCI MATLAB API.
Source code available at
https://github.com/token-cubed/UTK-ML-Research
____________________________________________________________________
Fisk University
Group: NSF—CREST (BioSS program)
Dates: Summer 2023
Mentors:
This NSF-funded internship was a deep dive into protein purification and crystallization (and my first introduction to academic research).
I was the lone CS student in a group of mostly biochemists.
Over the course
of the summer, I carried out many of the major steps of the protein purification and crystallization process, including:
____________________________________________________________________
|
||