Make runs easier to trust
Ran RNA-seq and TWAS across ~1,100 human brain cortex samples on SLURM, resolved Nextflow memory/input/batch issues, and built Python checks for run status.
Loading portfolio
Bioinformatics Engineer / Analyst
I’m Pranava. I turn messy sequencing projects into runnable pipelines, quality checks, and cautious biological readouts that collaborators can review.
Looking for roles where I can help teams move from raw sequencing data to checked outputs, clear reports, and well-scoped conclusions.
Evidence in:
The work has two sides: did the computation finish cleanly, and does the result support the story being told?
Ran RNA-seq and TWAS across ~1,100 human brain cortex samples on SLURM, resolved Nextflow memory/input/batch issues, and built Python checks for run status.
Across RNA-seq, spatial, single-cell, and ctDNA work, I separate observed signal from what the dataset cannot prove.
I write reports, dashboards, notebooks, and notes that leave the next person with the method, output, caveats, and next checks.
My lane is the handoff between computation and biology: get the data moving, check the output, and explain the result plainly.
At UT Dallas, I worked on RNA-seq and TWAS runs on SLURM, debugged Nextflow memory/input/batch failures, built Python status checks, and shared findings in reports and project meetings.
I am also interested in spatial transcriptomics methods development and PhD study, especially work that makes spatial results easier to validate and explain.
Earlier, I was closer to the bench: RNA-seq library prep, PCR, R/Bash scripts, antimicrobial assays, and SOP-driven microbiology. That background is why I care about both the system and the sample behind it.
Condensed resume evidence: HPC execution, R/Python scripting, quality reporting, and biology-facing documentation.
Ran RNA-seq and TWAS across ~1,100 human brain cortex samples on Linux SLURM HPC. Debugged three recurring Nextflow failure modes: memory limits, malformed inputs, and partial batches. Built Python checks with pandas and matplotlib, then communicated findings through reports and project meetings.
Prepared bulk RNA-seq libraries and PCR-amplified genomic targets, then analyzed expression data in R. Built reusable R and Bash scripts across four analysis steps: quality control, normalization, differential expression, and GSEA, with annotated R Markdown reporting.
University of Texas at Dallas, May 2025. Earlier training includes a B.S. in Microbiology, Biochemistry & Chemistry from Sri Venkateshwara University, plus microbiology experience in antimicrobial susceptibility testing and SOP documentation.
Look for command-line entry points, parser tests, summaries, model cards, and notes. Those details matter more than a polished figure alone.
The spatial, single-cell, and ctDNA projects include biological context, but the conclusions stay proportional to the data and validation available.
TinyVariant and run-risk show the habit I want to keep: start simple, check leakage, document limits, and explain the result even when the simpler method wins.
They show my preferred pattern: make the process observable, produce a concrete artifact, and document the boundary.
Built a Python CLI/package for nf-core/rnaseq triage. It reads trace/log telemetry, labels failed or resource-risky runs, trains interpretable logistic models, and writes JSON/Markdown reports.
With more real run history, I would test whether the risk labels stay useful outside small test-profile runs.
Open repository
Ported the public OpenVax neoantigen vaccine pipeline shape from Snakemake into a Nextflow DSL2 implementation, keeping the first pass close to the original command logic before modernization.
The useful part was keeping the biological logic conservative while making execution, resources, and outputs easier to review.
Open repository
Designed a seven-step Nextflow/R pipeline for public DLPFC Visium data: spot filtering, normalization, domain detection, SVG ranking, neighborhood summaries, and Shiny/Plotly review.
Next, I would add a true SpaceRanger ingest path only if raw 10x inputs are available.
Open repository
Built a reproducible ClinVar missense benchmark around a Tiny Recursive Model fork, then documented the important result: logistic regression remained stronger than the neural approach.
What surprised me was that the simpler baseline stayed stronger, which made the negative result more useful than a cleaner model story.
Open repository
Processed public colorectal cancer single-cell data through quality filtering, integration, clustering, marker-informed annotation, and tumor-normal composition visualization.
With more time, I would make the annotation step less manual and keep intermediate tables beside the figures.
Open repositoryNextflow, nf-core, Snakemake, CellRanger, SpaceRanger, Docker, Singularity, Conda, Git, GitHub Actions, pytest, and documentation for repeatable runs.
RNA-seq, TWAS, scRNA-seq, snRNA-seq, ATAC-seq, 10x Visium, FASTQ/BAM/VCF/H5AD, 10x outputs, GEO, SRA, and TCGA-style public datasets.
Scanpy, Seurat, scVI, SpatialExperiment, BayesSpace, Moran's I, DESeq2, edgeR, limma, Harmony, and CellChat.
scikit-learn, PyTorch, logistic baselines, ROC AUC, leakage checks, R Shiny, Plotly, dashboards, and reports.