Data Scientist · Drug Discovery
I work at the intersection of machine learning, molecular biology, and rigorous data engineering — turning complex biological signals into decisions that move drug discovery forward. From fine-tuning protein language models to curating analysis-ready datasets with full provenance, I care as much about how data is structured as what the model does with it.
01 — Projects
Machine Learning & AI
Protein ML · Immunology · NLP
ASPred — BCR Specificity via Protein Language Models
PEFT fine-tuned ESM2/ProtT5 for antigen-specific B Cell Receptor classification using contrastive learning. Careful repertoire-level data splits prevent leakage — a critical, non-obvious challenge.
Causal Discovery · Microbiome · Health
Causal Signatures in Gut Microbiome & Type 2 Diabetes
Applied PC, FCI, and GES causal discovery algorithms to identify disease-driving gut microbes beyond simple association. Interactive dashboard deployed live.
Data Analysis
Data Engineering · Drug Discovery · SQL
Scientific Dataset Curation for Drug Discovery ML
End-to-end Python/SQL pipeline on Linux for multi-source scientific dataset ingestion. SQL-backed metadata schemas track full provenance — enabling single-query audit trails for ML training data.
Clinical Data · -Omics · Reproducibility
Exploratory Biomarker Structures for Early Clinical Studies
Flexible, audit-ready dataset structures for non-standard endpoints in early-phase trials. Modular Python/R with Git ensures reproducible analysis from raw biomarker data to stakeholder reports.
Metagenomics · Cloud · Bioinformatics
Soil Microbiome Profiling at Loam Bio
Large-scale metagenomic pipeline (Qiime2, Kraken2, HUMAnN) containerized in Docker, deployed on AWS S3. Interactive Plotly/Dash dashboards for scientific and leadership stakeholders.
ML · Structural Biology · NIH-Funded
Antibody–Antigen Structural Energy Analysis
Random Forest and SVM models trained on Rosetta SnugDock docking simulations to identify intermolecular energy patterns predictive of high-affinity antibody binding.
Data Visualization
Dimensionality Reduction · Embeddings
Immune Repertoire Embedding Landscapes
Interactive UMAP/t-SNE visualizations of thousands of BCR embeddings from ESM2, revealing antigen-specific clustering not observable in raw sequence space. Validated that embedding geometry separates antigen classes, supporting downstream classification modeling.
Network Visualization · Dashboard
Microbiome Causal Network Dashboard
Interactive Plotly/Dash network graph of gut microbial causal relationships with filtering by edge confidence and taxonomy. Built for both scientific and non-specialist audiences.
02 — Stack
ML & Modeling — PyTorch, transformer fine-tuning, LoRA/PEFT, contrastive learning, ESM2, ProtT5, scikit-learn, distributed training on HPC/Slurm.
Data Engineering — Python, SQL, Linux/Bash, Pandas, NumPy, Git, Jupyter, Docker, Singularity, AWS/GCP, metadata cataloging, data provenance, audit-ready workflows, RESTful APIs.
Bioinformatics — Qiime2, Kraken2, HUMAnN, GROMACS, Rosetta/SnugDock, ProteinMPNN, AlphaFold/ESMFold, V(D)J pipelines.
Visualization & Communication — Plotly/Dash, ggplot2, UMAP, t-SNE, Seaborn. Fluent Spanish. Experienced communicating to non-technical stakeholders.
03 — Career Experience
04 — Contact
Open to roles in drug discovery data science, clinical data programming, and applied ML research. Always happy to connect.
marianapaco.m.21@gmail.com