Biochemistry, Department of
Department of Biochemistry: Dissertations, Theses, and Student Research
Accessibility Remediation
If you are unable to use this item in its current form due to accessibility barriers, you may request remediation through our remediation request form.
First Advisor
Tomáš Helikar
Second Advisor
Lindsey Crawford
Committee Members
Toshihiro Obata, Xinghui Sun, Massimiliano Pierobon
Date of this Version
7-2026
Document Type
Dissertation
Citation
A dissertation presented to the faculty of the Graduate College at the University of Nebraska in partial fulfillment of requirements for the degree of Doctor of Philosophy
Major: Biochemistry (Bioinformatics)
Under the supervision of Professors Tomáš Helikar and Lindsey Crawford
Lincoln, Nebraska, July 2026
Abstract
Cellular metabolism is reshaped by disease and physiological state, and genome-scale metabolic networks (GSMNs) offer a systematic way to interrogate this biology and identify therapeutic targets. Building a GSMN from expression data, however, requires a sequence of stages including read processing, normalization and gene-activity classification, multi-omic integration, model reconstruction, simulation, and biological interpretation. These are typically handled by disconnected tools spanning multiple programming languages, file formats, and databases. The resulting reliance on heterogeneous interfaces and manual steps threatens reproducibility. This work develops an ecosystem of computational tools that carry an analysis from raw sequencing reads to mechanistic biological interpretation. First, AutoRNAseq, a Snakemake-based bulk RNA-seq pipeline, alleviates manual, error-prone alignment pipelines by unifying data acquisition, reference genome preparation, quality control, alignment, and quantification. It was validated against a widely used reference pipeline (Pearson correlation > 98%). COMO consumes these outputs and integrates multi-omic processing, context-specific reconstruction, and drug-repurposing analysis in one workflow, only requiring parameter configuration. COMO was used to reconstruct B-cell metabolism in two disease contexts, and nominated ranked metabolic drug targets, several supported by clinical use or independent literature. A Python reimplementation of the R-based zFPKM gene-activity normalization method reproduced the original classifications (with Pearson and Spearman correlations both 1.00). Utilizing COMO at scale, a study of immune aging reconstructed 5,705 donor- and cell-type-specific models from a single-cell atlas of two million cells across 166 donors (aged 25-85), finding that age-associated metabolic signal concentrates in a small set of pathways and cell types, and shifts in a cell-type-dependent direction rather than uniformly. Lastly, MechAInistic, a tool-grounded, Architect-Reviewer large-language-model system, translates natural-language questions into executable, GSMN-oriented workflows that produce literature-supported hypotheses, outperforming general-purpose frontier models in grounding and task completion; MechAInistic provides an easy-to-use web interface and generates a literature-backed report, eliminating the need to program analysis pipelines. Together, these projects provides a reproducible, end-to-end ecosystem for expression-driven metabolic modeling, with each project solving a specific problem while reducing, or eliminating, the need for researchers to write code.
Advisors: Tomáš Helikar and Lindsey Crawford
Supplementary Information
Chapter 5 - Supplementary Tables.xlsx (324 kB)
Chapter 5 Supplementary Tables
Supplementary Figures.pdf (43303 kB)
Chapter 5 Supplementary Figures
Included in
Biochemistry Commons, Bioinformatics Commons, Medicinal and Pharmaceutical Chemistry Commons
Comments
Copyright 2026, Joshua Loecker. Used by permission