Browse Articles

Predicting smoking status based on RNA sequencing data

Yang et al. | Aug 30, 2024

Predicting smoking status based on RNA sequencing data
Image credit: Yang and Stanley 2024

Given an association between nicotine addiction and gene expression, we hypothesized that expression of genes commonly associated with smoking status would have variable expression between smokers and non-smokers. To test whether gene expression varies between smokers and non-smokers, we analyzed two publicly-available datasets that profiled RNA gene expression from brain (nucleus accumbens) and lung tissue taken from patients identified as smokers or non-smokers. We discovered statistically significant differences in expression of dozens of genes between smokers and non-smokers. To test whether gene expression can be used to predict whether a patient is a smoker or non-smoker, we used gene expression as the training data for a logistic regression or random forest classification model. The random forest classifier trained on lung tissue data showed the most robust results, with area under curve (AUC) values consistently between 0.82 and 0.93. Both models trained on nucleus accumbens data had poorer performance, with AUC values consistently between 0.65 and 0.7 when using random forest. These results suggest gene expression can be used to predict smoking status using traditional machine learning models. Additionally, based on our random forest model, we proposed KCNJ3 and TXLNGY as two candidate markers of smoking status. These findings, coupled with other genes identified in this study, present promising avenues for advancing applications related to the genetic foundation of smoking-related characteristics.

Read More...

Applying centrality analysis on a protein interaction network to predict colorectal cancer driver genes

Saha et al. | Nov 18, 2023

Applying centrality analysis on a protein interaction network to predict colorectal cancer driver genes

In this article the authors created an interaction map of proteins involved in colorectal cancer to look for driver vs. non-driver genes. That is they wanted to see if they could determine what genes are more likely to drive the development and progression in colorectal cancer and which are present in altered states but not necessarily driving disease progression.

Read More...

Integrated expression, mutation, and survival analysis of 17 key genes in breast cancer using TCGA-BRCA data

Wang et al. | Aug 06, 2026

Integrated expression, mutation, and survival analysis of 17 key genes in breast cancer using TCGA-BRCA data

This study combines gene expression, mutation profiling, and survival analysis of 17 clinically important genes in breast cancer, utilizing the TCGA-BRCA dataset. Our results show that there are different patterns of oncogene upregulation, different levels of tumor suppressor activity, and complicated survival associations. TP53 was the most frequently mutated gene in this cohort. The results underscore the significance of multidimensional genomic analyses for a comprehensive understanding of breast cancer biology and its therapeutic ramifications.

Read More...

Investigating the inhibition of catabolic enzymes for implications in cardiovascular diseases and diabetes

Gandhi et al. | Aug 25, 2024

Investigating the inhibition of catabolic enzymes for implications in cardiovascular diseases and diabetes
Image credit: The authors

Enzymes that metabolize carbohydrates and lipids play a key role in our health, including global health challenges like cardiovascular diseases and diabetes. To learn more about these important enzymes, Gandhi and Gandhi test whether various natural substances (ginger, Aloe vera, lemon, and mint leaves) affect the activity of α-amylase and lipase enzymes.

Read More...