Loading portfolio

Bioinformatics Engineer / Analyst

Omics pipelines that hold up after the first run.

I’m Pranava. I turn messy sequencing projects into runnable pipelines, quality checks, and cautious biological readouts that collaborators can review.

Looking for roles where I can help teams move from raw sequencing data to checked outputs, clear reports, and well-scoped conclusions.

Proof points ~1,100 samples 51 mapped rules 63,689 cells AUC 0.959

Evidence in:

RNA-seq/TWAS on SLURM OpenVax Nextflow port Python reporting Docker mini-runtime validation Nextflow failure triage Single-cell annotation Visium review ClinVar benchmarking
What I bring

I sit between the run log and the biological question.

The work has two sides: did the computation finish cleanly, and does the result support the story being told?

01

Make runs easier to trust

Ran RNA-seq and TWAS across ~1,100 human brain cortex samples on SLURM, resolved Nextflow memory/input/batch issues, and built Python checks for run status.

02

Keep interpretation grounded

Across RNA-seq, spatial, single-cell, and ctDNA work, I separate observed signal from what the dataset cannot prove.

03

Package the handoff

I write reports, dashboards, notebooks, and notes that leave the next person with the method, output, caveats, and next checks.

Meet

Pranava

Pranava Upparlapalli headshot
M.S. Bioinformatics & Computational Biology, UT Dallas.

My lane is the handoff between computation and biology: get the data moving, check the output, and explain the result plainly.

At UT Dallas, I worked on RNA-seq and TWAS runs on SLURM, debugged Nextflow memory/input/batch failures, built Python status checks, and shared findings in reports and project meetings.

I am also interested in spatial transcriptomics methods development and PhD study, especially work that makes spatial results easier to validate and explain.

Earlier, I was closer to the bench: RNA-seq library prep, PCR, R/Bash scripts, antimicrobial assays, and SOP-driven microbiology. That background is why I care about both the system and the sample behind it.

Experience

The work history behind the projects.

Condensed resume evidence: HPC execution, R/Python scripting, quality reporting, and biology-facing documentation.

Jan 2025 - May 2025

Bioinformatics Analyst, University of Texas at Dallas

Ran RNA-seq and TWAS across ~1,100 human brain cortex samples on Linux SLURM HPC. Debugged three recurring Nextflow failure modes: memory limits, malformed inputs, and partial batches. Built Python checks with pandas and matplotlib, then communicated findings through reports and project meetings.

  • RNA-seq / TWAS
  • SLURM HPC
  • Nextflow debugging
  • Python status checks
Jun 2022 - May 2023

Bioinformatics Research Associate, Bharati Vidyapeeth University

Prepared bulk RNA-seq libraries and PCR-amplified genomic targets, then analyzed expression data in R. Built reusable R and Bash scripts across four analysis steps: quality control, normalization, differential expression, and GSEA, with annotated R Markdown reporting.

  • Bulk RNA-seq
  • R / Bash
  • DE analysis
  • R Markdown
Education

M.S. Bioinformatics & Computational Biology

University of Texas at Dallas, May 2025. Earlier training includes a B.S. in Microbiology, Biochemistry & Chemistry from Sri Venkateshwara University, plus microbiology experience in antimicrobial susceptibility testing and SOP documentation.

  • UT Dallas
  • Microbiology foundation
  • CLSI-style assays
  • Reproducible records
Review guide

How to judge the work.

01

Can it be rerun?

Look for command-line entry points, parser tests, summaries, model cards, and notes. Those details matter more than a polished figure alone.

02

Are the claims bounded?

The spatial, single-cell, and ctDNA projects include biological context, but the conclusions stay proportional to the data and validation available.

03

Are baselines visible?

TinyVariant and run-risk show the habit I want to keep: start simple, check leakage, document limits, and explain the result even when the simpler method wins.

Selected work

Start with these projects.

They show my preferred pattern: make the process observable, produce a concrete artifact, and document the boundary.

01
nf-core rnaseq run-risk workflow diagram
Pipeline reliability / Python CLI

nf-core/rnaseq Run-Risk Predictor

Built a Python CLI/package for nf-core/rnaseq triage. It reads trace/log telemetry, labels failed or resource-risky runs, trains interpretable logistic models, and writes JSON/Markdown reports.

  • CLI package
  • Parser tests
  • Model card
  • Risk reports
Proof
make-runs, parse, train, predict, mega-run commands
Boundary
Proof-of-concept scope on 51 small test-profile runs
Evidence
Telemetry parsing, null handling, tests, model card, and reports

With more real run history, I would test whether the risk labels stay useful outside small test-profile runs.

Open repository
02
OpenVax Nextflow v1.2 workflow diagram
Nextflow port / workflow engineering

OpenVax Nextflow Port

Ported the public OpenVax neoantigen vaccine pipeline shape from Snakemake into a Nextflow DSL2 implementation, keeping the first pass close to the original command logic before modernization.

  • 51/51 rules mapped
  • Docker image
  • Patient manifests
  • Mini runtime
Proof
Connected faithful_port workflow, startup validation, run metadata, reports, and manifests
Runtime
Two-patient Docker mini run with 106 succeeded Nextflow tasks
Boundary
Rule-level parity and local Docker validation, not clinical or cloud runtime validation

The useful part was keeping the biological logic conservative while making execution, resources, and outputs easier to review.

Open repository
03
Spatial transcriptomics workflow preview
Spatial transcriptomics / Nextflow / R

Spatial Insights Cohort

Designed a seven-step Nextflow/R pipeline for public DLPFC Visium data: spot filtering, normalization, domain detection, SVG ranking, neighborhood summaries, and Shiny/Plotly review.

  • Nextflow/R
  • Spot filters
  • BayesSpace
  • Shiny/Plotly
Scale
2 Visium samples, 7,783 spots
Boundary
Pipeline and dashboard engineering with descriptive interpretation
Artifacts
Validation report, step summaries, per-sample outputs, and spatial quality artifacts

Next, I would add a true SpaceRanger ingest path only if raw 10x inputs are available.

Open repository
04
TinyVariant benchmark preview
ML evaluation / ClinVar

TinyVariant Benchmark

Built a reproducible ClinVar missense benchmark around a Tiny Recursive Model fork, then documented the important result: logistic regression remained stronger than the neural approach.

  • Leakage checks
  • Baseline first
  • Sweep tooling
  • AUC report
Result
Logistic regression AUC 0.959 vs best TRM AUC 0.951
Signal
Evaluation discipline over model hype
Checks
Preprocessing, leakage tests, ablations, sweep summary, and interpretation

What surprised me was that the simpler baseline stayed stronger, which made the negative result more useful than a cleaner model story.

Open repository
05
Colorectal cancer single-cell analysis preview
Single-cell analysis

CRC TME scRNA-seq

Processed public colorectal cancer single-cell data through quality filtering, integration, clustering, marker-informed annotation, and tumor-normal composition visualization.

  • 63,689 cells
  • Scanpy/scVI
  • UMAPs

With more time, I would make the annotation step less manual and keep intermediate tables beside the figures.

Open repository
Stack

The stack behind the work.

01

Pipeline systems

Nextflow, nf-core, Snakemake, CellRanger, SpaceRanger, Docker, Singularity, Conda, Git, GitHub Actions, pytest, and documentation for repeatable runs.

02

Sequencing data

RNA-seq, TWAS, scRNA-seq, snRNA-seq, ATAC-seq, 10x Visium, FASTQ/BAM/VCF/H5AD, 10x outputs, GEO, SRA, and TCGA-style public datasets.

03

Analysis libraries

Scanpy, Seurat, scVI, SpatialExperiment, BayesSpace, Moran's I, DESeq2, edgeR, limma, Harmony, and CellChat.

04

Evaluation and reporting

scikit-learn, PyTorch, logistic baselines, ROC AUC, leakage checks, R Shiny, Plotly, dashboards, and reports.

Contact

Open to roles where pipelines and omics interpretation meet.