Data Scientist · Drug Discovery

Building data
pipelines that
science can trust.

I work at the intersection of machine learning, molecular biology, and rigorous data engineering — turning complex biological signals into decisions that move drug discovery forward. From fine-tuning protein language models to curating analysis-ready datasets with full provenance, I care as much about how data is structured as what the model does with it.

Location USA
Education BS Data Science, UC San Diego
Published bioRxiv · NeurIPS 2024
Email marianapaco.m.21@gmail.com

01 — Projects

Selected work

Machine Learning & AI

Protein ML · Immunology · NLP

ASPred — BCR Specificity via Protein Language Models

PEFT fine-tuned ESM2/ProtT5 for antigen-specific B Cell Receptor classification using contrastive learning. Careful repertoire-level data splits prevent leakage — a critical, non-obvious challenge.

NeurIPS 2024 bioRxiv

Causal Discovery · Microbiome · Health

Causal Signatures in Gut Microbiome & Type 2 Diabetes

Applied PC, FCI, and GES causal discovery algorithms to identify disease-driving gut microbes beyond simple association. Interactive dashboard deployed live.

Live

Data Analysis

Data Engineering · Drug Discovery · SQL

Scientific Dataset Curation for Drug Discovery ML

End-to-end Python/SQL pipeline on Linux for multi-source scientific dataset ingestion. SQL-backed metadata schemas track full provenance — enabling single-query audit trails for ML training data.

Clinical Data · -Omics · Reproducibility

Exploratory Biomarker Structures for Early Clinical Studies

Flexible, audit-ready dataset structures for non-standard endpoints in early-phase trials. Modular Python/R with Git ensures reproducible analysis from raw biomarker data to stakeholder reports.

Metagenomics · Cloud · Bioinformatics

Soil Microbiome Profiling at Loam Bio

Large-scale metagenomic pipeline (Qiime2, Kraken2, HUMAnN) containerized in Docker, deployed on AWS S3. Interactive Plotly/Dash dashboards for scientific and leadership stakeholders.

ML · Structural Biology · NIH-Funded

Antibody–Antigen Structural Energy Analysis

Random Forest and SVM models trained on Rosetta SnugDock docking simulations to identify intermolecular energy patterns predictive of high-affinity antibody binding.

Data Visualization

Dimensionality Reduction · Embeddings

Immune Repertoire Embedding Landscapes

Interactive UMAP/t-SNE visualizations of thousands of BCR embeddings from ESM2, revealing antigen-specific clustering not observable in raw sequence space. Validated that embedding geometry separates antigen classes, supporting downstream classification modeling.

Network Visualization · Dashboard

Microbiome Causal Network Dashboard

Interactive Plotly/Dash network graph of gut microbial causal relationships with filtering by edge confidence and taxonomy. Built for both scientific and non-specialist audiences.

Live

02 — Stack

Technical capabilities

ML & Modeling — PyTorch, transformer fine-tuning, LoRA/PEFT, contrastive learning, ESM2, ProtT5, scikit-learn, distributed training on HPC/Slurm.

Data Engineering — Python, SQL, Linux/Bash, Pandas, NumPy, Git, Jupyter, Docker, Singularity, AWS/GCP, metadata cataloging, data provenance, audit-ready workflows, RESTful APIs.

Bioinformatics — Qiime2, Kraken2, HUMAnN, GROMACS, Rosetta/SnugDock, ProteinMPNN, AlphaFold/ESMFold, V(D)J pipelines.

Visualization & Communication — Plotly/Dash, ggplot2, UMAP, t-SNE, Seaborn. Fluent Spanish. Experienced communicating to non-technical stakeholders.

03 — Career Experience

04 — Contact

Let's work
together.

Open to roles in drug discovery data science, clinical data programming, and applied ML research. Always happy to connect.

marianapaco.m.21@gmail.com