THE PROFILE

Trevor Santiago

Applied AI engineer

I build AI applications that turn messy data into useful tools.

Sunlit stone, ferns, and timber in the 16-bit Garden hall. A little world to explore

Explore the work in the Profile or Game. AI Q&A is under construction.

ABOUT

A bit of context.

I'm Trevor Santiago, a data scientist at ALLDATA building AI applications and data workflows. My recent work includes Gemini-powered services, semantic retrieval, and tools that help subject-matter experts reuse prior edits and review AI-generated drafts.

I like taking an unclear problem through research, prototyping, and implementation. At ALLDATA, that has meant working directly with subject-matter experts, building FastAPI services on Google Cloud, and checking model output against the cases those experts care about.

Outside work, I built NormFlow, a text-normalization workbench that combines exact matching, semantic retrieval, and optional LLM suggestions with batch processing and human review.

Sacramento, California · English, fluent/native

US citizen, no sponsorship required.

US remote work preferred. Open to relocation for hybrid work in the San Francisco Bay Area or San Diego.

EXPERIENCE

The work behind the work.

November 2021 to present

Data Scientist

ALLDATA · Elk Grove, California

Build AI services, human review workflows, and cloud data pipelines.

  • Built the Repair Trends AI service with FastAPI, Cloud Run, Gemini through Vertex AI, and an AlloyDB cache that regenerates results when upstream component inputs change. The beta/staging rollout is complete.
  • Independently built a Probable Cause Data editing system that fills text edits from a library of previously applied edits using progressive matching. When no match is found, it uses a separately fine-tuned Flan-T5 Small model to generate a new draft edit.
  • Fine-tuned the generation model with QLoRA and evaluated it with BLEU separately from the editing workflow. The primary SME estimated that the overall system reduced total editing effort by at least 50%.
  • Owned the Known Fixes pipeline for billions of dealership repair-order records, using Gemini with web-search grounding to standardize free-text part descriptions before ACES mapping.
  • Combined repair-order data with SME-curated Probable Causes, applied related-vehicle rules, and automated monthly Known Fixes publishing with BigQuery, Cloud Run Jobs, Airflow, and AlloyDB. Known Fixes is a core feature of Diagnostic Intelligence, which reached 1,700+ subscriptions in its first three months.

January 2021 to August 2021

Data Science Intern

New York Mets

Built predictive-model components and venue features for a larger outfield-alignment system.

  • Owned an XGBoost hit-outcome classifier within a multi-model system designed to maximize expected outs for batter and pitcher matchups.
  • Trained and evaluated the classifier using separate historical training, evaluation, and held-out test sets.
  • Standardized wall geometry for every MLB venue with interpolation and spline fitting, estimating distance from home plate at one-degree increments.

FEATURED PROJECTS

Things built, taken apart,
and built again.

Independent project

A text-normalization workbench with exact matching, semantic retrieval, optional LLM fallback, and human review.

  • Python
  • FastAPI
  • TypeScript
  • SQLite
  • FAISS
  • Semantic retrieval
  • Human review
Read project details

ALLDATA

An editing workflow that fills from previously applied edits, then uses a separately fine-tuned model to generate new draft edits when matching finds no reusable edit.

  • Python
  • Semantic retrieval
  • Flan-T5
  • QLoRA
  • BLEU
  • Cloud Run Jobs
  • Cloud Storage
Read project details

PROJECT ARCHIVE

More from the workbench.

Known Fixes

ALLDATA

A repair-intelligence pipeline combining dealership repair orders and expert-curated probable causes to rank parts associated with diagnostic trouble codes.

Find a Fix data pipeline

ALLDATA · Retired project

BigQuery and Dataform pipelines over millions of retail-sales, lookup, and scan records for a legacy repair-ranking product.

SKILLS & EDUCATION

Tools and foundations.

  • Python
  • SQL
  • LLM applications
  • Semantic retrieval
  • Prompt engineering
  • Data pipelines
  • FastAPI
  • Flask
  • Docker
  • TypeScript
  • SQLite
  • scikit-learn
  • XGBoost
  • Pandas
  • NumPy
  • FAISS
  • AlloyDB / PostgreSQL
  • BigQuery
  • Dataform
  • Airflow
  • Vertex AI
  • Cloud Run
  • Cloud Functions
  • Cloud Storage
  • Solr
  • Claude Code
  • Codex
  • Vercel

August 2020 to August 2021

M.S. in Data Science

University of San Francisco

Selected coursework: machine learning, deep learning, databases, distributed computing, and design of experiments.

July 2016 to June 2020

B.S. in Mathematical Sciences

University of California, Santa Barbara

Minor in Statistical Science. Selected coursework: linear algebra, probability, stochastic processes, operations research, and linear regression.

Certifications

  • Fundamentals of LLMsHugging Face · June 2025
  • LangChain Chat with Your DataDeepLearning.AI · May 2025
  • LangChain for LLM Application DevelopmentDeepLearning.AI · May 2025
  • Free Data Engineering Bootcamp CertificateDataExpert.io · January 2025
  • Google Cloud Fundamentals: Core InfrastructureGoogle · November 2023

Recognition

  • Extra Miler of the YearALLDATA · 2023

AI Q&A

Ask about the work.

UNDER CONSTRUCTION

AI Q&A is on the way.

I'm building and testing a Q&A that answers questions about my work and links to the supporting evidence.

For now, you can read the projects and experience directly or get in touch.

LINKS & CONTACT

The usual places.

trevorjsantiago1@gmail.com