Skip to content
All projects
03Case study

Financial Market Sentiment Classification

A deep-learning NLP system that classifies the dominant sentiment of financial news. Custom embedding implementation with NumPy and neural-network training using PyTorch/CUDA, with MLP and self-attention components.

Context
Personal project
Role
Design & implementation
Timeline
2026/05/01
Stack
Python · NumPy · PyTorch · CUDA
GitHubSoon

01 — Overview

Overview

A neural NLP classifier that reads financial news and predicts its dominant sentiment. The embedding layer is a custom implementation written with NumPy; the neural network is trained with PyTorch on GPU through CUDA.

02 — Problem

Problem

Financial news is a high-volume input for market analysis. Classifying its dominant sentiment is a first step towards using text as a structured signal.

the Sentiments classes are Positive , negative , neutral . I choosed to implement some tools from scratch to understand deeply the underlying mechanisms

03 — Architecture

Architecture

  1. Financial news text
  2. Tokenization
  3. EmbeddingsCustom implementation · NumPy
  4. Self-attention
  5. MLP · hidden layers
  6. Softmax classification
  7. Dominant sentiment

Conceptual pipeline .

04 — Data / Inputs

Data / Inputs

Dataset source was Bloomberg news articles from hugging-face , size : 144000 articles , classes in the training corpus was clearly balanced 37% positive , 40% negative , 23% neutral

05 — Methodology

Methodology

  1. 01TokenizationRaw news text is split into tokens and mapped to vocabulary indices.
  2. 02EmbeddingsCustom embedding implementation with NumPy.
  3. 03Self-attentionLets each token representation weigh the rest of the sequence.
  4. 04MLPHidden layers transform the attended representation.
  5. 05SoftmaxProduces a probability distribution over sentiment classes.

Training uses backpropagation with the Adam optimizer and mini-batch updates, accelerated on GPU with PyTorch and CUDA.

  • Backpropagation
  • Adam optimizer
  • Mini-batch training
  • Softmax
  • Self-attention
  • MLP

06 — Engineering Implementation

Engineering Implementation

  • NumPy for the custom embedding implementation.
  • PyTorch for the neural network and training loop.
  • CUDA for GPU-accelerated training.

07 — Evaluation / Results

Evaluation / Results

MetricValidationTest
Accuracy62 % 67 %
Macro F166 % 59 %
Terminal output: tokenizer training with a 30,000-word vocabulary, then 10 training epochs with train/validation loss, accuracy and F1, and final test results (accuracy 0.66, macro F1 0.59)
WordPiece tokenizer training (30,000-token vocabulary) and 10-epoch model training on an RTX 3080 (CUDA)
Terminal output: per-class precision, recall and F1 for Negative, Neutral and Positive, the confusion matrix, and a new article classified as Neutral with 0.99 probability
Classification report, confusion matrix and inference on an unseen article

08 — Challenges & Trade-offs

Challenges & Trade-offs

Challenges faced are transformation of all articles into numerical representations and handling class imbalance.

09 — Resources

Resources

https://github.com/charfx/NLP_financial_sentiment