Visual-Valence Model (VCA)

A deep neural network model of associative emotional (Pavlovian fear) learning.

Companion model repository for "Associative Emotional Learning in Convolutional Neural Networks" (Leem, Keil, Ding & Fang; Neural Computation, in press).

Note: This model is a research artifact for computational/cognitive neuroscience, released to reproduce and extend the paper's findings. It is not a general-purpose sentiment/emotion API and is not validated for clinical, diagnostic, or production affective-computing use.


Model description

The Visual-Valence Model predicts the affective valence of a visual scene (1 = extreme displeasure, 9 = extreme pleasure) and reproduces hallmarks of Pavlovian (fear) conditioning when a novel, initially neutral stimulus is repeatedly paired with an emotionally-charged one.

The architecture (Visual_Cortex_Amygdala in models/VGG_Model.py) has three components, loosely modeling the primate visual/affective pathway:

Component Role Implementation
Visual Cortex module ("High Road") Ventral visual stream VGG-16 (Simonyan & Zisserman, 2015), ImageNet-pretrained, frozen
Shortcut Pathway ("Middle Road") Fast, low-resolution route from early vision to affect circuitry Multi-scale pooling of VGG-16's early layers (layer index 10), combined via an Efficient Channel Attention (ECA) module (Wang et al., 2020)
Valence Module Combines both pathways into a scalar valence judgment Two fully-connected layers (amygdala LA/CE-nuclei- and OFC-inspired) + one sigmoid output unit

The High Road output (4096-d) and Middle Road output (512-d) are concatenated (4608-d) and passed through the Valence Module's fully connected layers to a single sigmoid unit, which is linearly rescaled from [0, 1] to the [1, 9] IAPS valence scale at evaluation time.

Five checkpoints, one training pipeline

This repository hosts every checkpoint along the training pipeline described in the GitHub README — same architecture (Visual_Cortex_Amygdala) throughout, so any of them can be loaded with the same code:

File Stage Description
vca_ckvideo_batch128_lr2e-5_epoch20.pth 0 Trained from scratch on the Cowen & Keltner Videoframe dataset
vca_IAPS_batch10_lr2e-4_epoch23.pth 1 Fine-tuned on full-size IAPS images
base_model_vca_IAPS_quadrant.pth 2 Fine-tuned to the quadrant-cropped input layout — pre-conditioning: has learned to decode valence from natural scenes (the US) but has never seen the conditioned stimulus (CS, a Gabor patch)
base_model_conditioned_orientation_epoch1.pth 3, epoch 1 First epoch of Pavlovian conditioning (CS+ Gabor patch paired with pleasant/unpleasant IAPS US); early/under-trained, kept for provenance
base_model_conditioned_orientation_epoch100.pth 3, epoch 100 Post-conditioning (final) — used throughout the paper's conditioning/generalization/representational-alignment analyses

Comparing the pre- (Stage 2) and post-conditioning (Stage 3, epoch 100) checkpoints' responses to the CS alone is what reveals the learned CS→valence association (and, at the representation level, the increasing CS/US alignment reported in the paper).


Intended use

  • Reproducing the paper's Pavlovian-conditioning experiments and figures.
  • Extending the model to new conditioning paradigms, stimulus sets, or ablations (e.g. the attention-free variant, Visual_Cortex_Amygdala_wo_Attention) for research purposes.
  • Studying how a two-pathway (ventral-stream + shortcut) architecture with attention reproduces behavioral/representational signatures of associative emotional learning.
  • Feature extraction / representational analysis (see Channel_Activity_Extraction.py, Manifold_Visualization.py, SVM_Analysis_*.py in the GitHub repo) for downstream neuroscience analyses.

Out of scope: general-purpose sentiment analysis, affect recognition on non-IAPS-like natural images, clinical/diagnostic use, or any decision-making about real individuals' emotional states.


Limitations

  • Frozen ImageNet backbone. The High Road (VGG-16) is never fine-tuned, so the model inherits ImageNet's visual biases; only the shortcut pathway and valence module are trained on affective data.
  • Narrow, licensed training data. IAPS is a relatively small, curated stimulus set under a data-use/confidentiality agreement (not redistributed with the code or this checkpoint); the model's generalization to naturalistic, in-the-wild imagery is untested.
  • Two-alternative conditioning paradigm. The conditioning stage pairs exactly two Gabor orientations (45°, 135°) with unpleasant/pleasant IAPS images; behavior on other CS types has not been evaluated.
  • Not validated against individual human subjects. Comparisons to human data in the paper are at the group/aggregate level; the model is not a predictive model of any specific person's affective response.
  • Single scalar output. The model predicts only valence (pleasant ↔ unpleasant), not arousal or discrete emotion categories.

Evaluation metrics

Model quality is reported using:

  • Pearson correlation (R / R²) between predicted and ground-truth (SAM-rated) valence, computed by reg_eval_model / cond_eval_model in utils.py.
  • Mean-squared error (MSE) between predicted and ground-truth valence (rescaled to the 1–9 scale).
  • Post-conditioning, generalization and CS/US representational alignment (single-unit tuning and population-level RSA/t-SNE, via Channel_Activity_Extraction.py, Manifold_Visualization.py, SVM_Analysis_Emotion.py, SVM_Analysis_Before_After.py) are used to assess whether conditioning reproduces human associative-learning signatures.

The full quantitative results (per-stage R/MSE, generalization curves, and alignment statistics) are reported in the paper's tables/figures — see https://arxiv.org/abs/2607.19327.


Training dataset

Training proceeds in stages, each building on the previous stage's checkpoint (see the GitHub README's Training section for exact commands):

Stage Dataset Purpose
0 Cowen & Keltner (2017) Videoframe dataset (2,185 emotion-eliciting video clips, one frame sampled every 10th frame) Pretrain valence regression from scratch on natural scenes
1 International Affective Picture System (IAPS; Bradley & Lang, 1994/2007), full-size images, 8:1:1 train/val/test split Adapt to the US stimuli used in conditioning
2 IAPS, quadrant-cropped layout Fine-tune to the spatial layout used during conditioning
3 IAPS (US, in the 4th quadrant) × Gabor patches (CS, in the 2nd quadrant; 45°/135° orientation, varying spatial frequency/contrast, generated via Gabor4Seowung.m) Pavlovian conditioning: pair CS+ with pleasant/unpleasant US

Labels are human valence ratings on a 1–9 scale (Self-Assessment Manikin; Bradley & Lang, 1994). The IAPS images themselves are not redistributed with the code or this model repository due to a data-use/confidentiality agreement — obtain access to IAPS independently to reproduce training from scratch. The Gabor-patch CS stimuli are procedurally generated and have no such restriction.

Note the input layout differs by stage: Stage 0/1 checkpoints expect a full-frame natural image (resize + normalize only), while Stage 2/3 checkpoints expect the quadrant-cropped layout (see Data preprocessing in the GitHub README) — feeding a full-frame image to a Stage 2/3 checkpoint (or vice versa) will not reproduce the reported behavior.


How to use

This is a plain PyTorch checkpoint (not a transformers model), so inference requires the model class definition from the companion GitHub repository. See inference_example.py in this repository for a complete, runnable example. In short:

git clone https://github.com/lab-smile/FearConditioningAI.git
cd FearConditioningAI
pip install -r requirements.txt   # or: conda env create -f environment-<platform>.yml
import torch
from huggingface_hub import hf_hub_download
from models.VGG_Model import Visual_Cortex_Amygdala

repo_id = "smilelab/visual-valence-model"
# swap in any filename from the checkpoint table above, e.g. "vca_ckvideo_batch128_lr2e-5_epoch20.pth"
ckpt_path = hf_hub_download(repo_id=repo_id, filename="base_model_conditioned_orientation_epoch100.pth")

model = Visual_Cortex_Amygdala()
checkpoint = torch.load(ckpt_path, map_location="cpu", weights_only=False)
model.load_state_dict(checkpoint["state_dict"], strict=False)
model.eval()

See inference_example.py for image preprocessing (resize/normalize + quadrant placement of the CS/US) and how to rescale the model's sigmoid output back to the 1–9 valence scale.


Citation

If you use this model, please cite the paper:

@article{leem2026associative,
  title   = {Associative Emotional Learning in Convolutional Neural Networks},
  author  = {Leem, Seowung and Keil, Andreas and Ding, Mingzhou and Fang, Ruogu},
  journal = {Neural Computation},
  year    = {2026},
  note    = {in press},
  eprint  = {2607.19327},
  archivePrefix = {arXiv},
  url     = {https://arxiv.org/abs/2607.19327}
}

Please also cite the datasets and methods this model builds on (IAPS, Cowen & Keltner Videoframe, SAM, VGG-16, ImageNet, ECA-Net, Rescorla-Wagner) — full references in the GitHub README's Citations section.


License

This model is released under the MIT License, matching the GitHub repository.

Contact

Name Email
Seowung Leem leem.s@ufl.edu
Dr. Ruogu Fang ruogu.fang@bme.ufl.edu
Downloads last month
30
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for smilelab/visual-valence-model