| dc.description.abstract | Brain tumors are among the most serious neurological conditions, where accurate
and timely diagnosis directly affects treatment planning and patient survival. Mag-
netic Resonance Imaging (MRI) is the standard modality for intracranial lesion
assessment, but manual interpretation is time-consuming, subject to inter-observer
variability, and demands specialized expertise. This thesis presents a unified deep
learning framework that performs brain tumor classification and segmentation from
2D MRI slices in a single forward pass. Although in deep learning the segmentation
step inherently produces a per-pixel class assignment, this work deliberately treats
classification (assigning a single tumor type to the whole scan) and segmentation
(delineating the tumor region) as two reported outputs of the same network, because
each output answers a different clinical question and is evaluated with a different
family of metrics. Three gaps in the literature are addressed: pronounced class
imbalance in public datasets, the limited ability of the standard U-Net to produce
fine tumor boundaries, and the architectural disconnect between classification and
segmentation in U-Net-based brain tumor pipelines. Two encoder–decoder networks
were developed and compared: a Standard U-Net baseline and an Enhanced U-
Net that integrates residual connections, attention gates at every skip junction, and
Squeeze-and-Excitation (SE) channel-recalibration modules. A composite training
objective combining categorical cross-entropy, focal loss, Dice loss, and a boundary-
aware term was used to jointly optimize overlap and contour accuracy. The training
image dataset was assembled by merging two publicly available repositories—the
Brain Tumor Segmentation Dataset by Akter et al. and the Brain Tumor Dataset
(MAT Format) by Cheng et al.—followed by quality control and hybrid resampling
to produce a class-balanced collection of 6,380 images with 1,595 samples per class
across four categories: no tumor, glioma, meningioma, and pituitary. On the held-out
test set, the Enhanced U-Net achieved pixel-wise accuracy of 99.73%, a Dice coeffi-
cient of 0.9141, and an IoU of 0.9947, improving on the Standard U-Net baseline
(99.62% accuracy, 0.8423 Dice, 0.9913 IoU). Single-image inference was measured
at 80–100 ms for the Standard U-Net and 100–150 ms for the Enhanced U-Net on an
NVIDIA Tesla P100 GPU, well below clinical real-time requirements. Image-level
classification, derived from spatial aggregation of segmentation probabilities, yielded
mean precision and recall of 0.983 (per-class values ranging from 0.94 to 1.00).
Comparative experiments on the two original imbalanced datasets confirmed the
importance of dataset curation: models trained on Dataset 1 achieved only 69% recall
on the minority glioma class, while models trained on Dataset 2 (which contains no
healthy scans) failed completely to recognise healthy MRI slices as no-tumor cases.
The ablation study confirmed that each architectural module contributes measurably,
and the three mechanisms (residual connections, attention gates, and SE modules)
reinforce one another. The proposed multi-task Enhanced U-Net pipeline—which
jointly produces a pixel-level segmentation map and an image-level class label in a
single forward pass—demonstrates that combining targeted architectural refinements
with a class-balanced dataset and a multi-task formulation produces a reliable tool for
computer-aided brain tumor analysis suitable as a basis for further clinical validation. | en_US |