Neurological disorders can alter motor control, cognition, and speech production in subtle ways that are difficult to detect using standard diagnostic tools. Because conventional imaging techniques such as CT and MRI often fail to reveal functional abnormalities, there is growing interest in non-invasive digital health approaches. This study investigates the potential of speech-based deep learning for detecting concussion-related dysfunction using structured speech tasks. Audio recordings from 225 high school and college athletes (approximately 200 retained after quality filtering) performing eight speech tasks were analyzed. The recordings were converted into Mel spectrograms and processed using a convolutional neural network (CNN), with data augmentation applied to improve model generalization. The results demonstrate strong classification performance, particularly for tasks requiring dynamic articulation and rhythmic control. To better interpret model behavior, the speech tasks were grouped into three functional categories and examined using gradient-weighted class activation mapping (Grad-CAM). Sentence and word reading tasks, as well as rapid syllable repetition, achieved the highest accuracy (up to 96%) and exhibited focused activation patterns associated with phonetic complexity and motor timing. In contrast, sustained vowel tasks produced lower performance and more diffuse model attention, likely due to their limited acoustic variability. These findings highlight the effectiveness of CNN-based analysis of dynamic speech features and emphasize the importance of task selection in speech-based digital diagnostics. Although the framework may inform future speech-based neurological assessment research, this study focuses on mild traumatic brain injury (concussion) as an initial use case.