Electrical and Computer Engineering, Department of

 

Department of Electrical and Computer Engineering: Dissertations, Theses, and Student Research

Accessibility Remediation

If you are unable to use this item in its current form due to accessibility barriers, you may request remediation through our remediation request form.

First Advisor

Hasan Otu

Committee Members

Mohammad Rashedul Hasan, Andrew Harms

Date of this Version

4-21-2026

Document Type

Thesis

Citation

A thesis presented to the Faculty of the Graduate College at the University of Nebraska in partial fulfillment of requirements for the Degree of Master of Science

Major: Electrical Engineering

Under the supervision of Professor Hasan Otu

Lincoln, Nebraska, April 2026

Comments

Copyright 2026, Cooper Schmer. Used by permission

Abstract

High-dimensional biomedical datasets, such as omics data, present significant challenges for predictive modeling due to noise, redundancy, and computational complexity. This thesis proposes a hybrid framework that integrates Bayesian Networks (BNs) and Artificial Neural Networks (NNs) to improve classification performance of such data sets while reducing input dimensionality. Central to this work is a novel feature selection method based on d-separation, a structural property of Bayesian networks that encodes conditional independence relationships.

The proposed approach introduces a count-based d-separation metric to quantify the relevance of variables to a target outcome, along with a thresholding scheme to balance feature selection robustness and sparsity. This method is evaluated against traditional approaches, including neural networks trained on all variables and those using Markov Blanket (MB)-based feature selection.

Experiments were conducted on both synthetic datasets and real-world cancer datasets from The Cancer Genome Atlas (TCGA), including pancreatic, colorectal, liver, and stomach cancers. Results demonstrate that the d-separation-based method consistently produces compact feature sets that retain essential predictive information. Across datasets, neural networks trained on d-separation-selected features achieved performance comparable to or better than models trained on the full feature set, while using significantly fewer variables. In contrast, MB-based methods exhibited instability in cases where network structure was sparse or degenerate.

To further enhance predictive performance, an ensemble framework combining BN and NN predictions was developed. While Bayesian networks alone showed weaker predictive power, the ensemble approach improved overall model stability and, in some cases, achieved performance gains over individual models.

Overall, this work demonstrates that incorporating structural information from probabilistic graphical models into neural network pipelines provides an effective and interpretable strategy for feature selection in high-dimensional biomedical data. The proposed framework offers a practical balance between model performance, robustness, and interpretability, with potential applications in precision medicine and multi-omics data analysis.

Advisor: Hasan Otu

Share

COinS