From Behavior to Pixels: Android Ransomware Detection with a Vision Transformer

Our ICDAM 2025 paper: sandbox behaviour of 4,280 Android apps turned into features and images, then classified by a Random Forest, CNNs and a ViT (99.78% accuracy).

Satyam KesharwaniSatyam Kesharwani 6 min read
  • Machine Learning
  • Security
  • Research

TL;DR. In our ICDAM 2025 paper, From Behavior to Pixels: A Vision Transformer Approach for Android Ransomware Detection (Springer LNNS, where I am the lead and corresponding author), we ran 4,280 Android apps in the CuckooDroid sandbox, summarised each app's behaviour as nine features, and classified it four ways: a Random Forest on the features, CNNs on RGB and grayscale images of the features, and a pretrained Vision Transformer (ViT). The ViT scored highest, at 99.78% accuracy (99.76% precision, recall and F1), followed by the RGB CNN at 99.76%, the grayscale CNN at 99.53% and the Random Forest at 99.41%. To the best of our knowledge, it was the first use of a ViT exclusively for Android ransomware detection. This post explains the pipeline, the "behaviour to pixels" idea, and how I read those numbers.

Paper: Springer · Code: GitHub

Why behaviour, and why Android ransomware

Signature-based antivirus matches known byte patterns, so a repackaged or slightly modified ransomware sample can slip past it. What ransomware cannot easily hide is what it does when it runs: requesting dangerous permissions, dropping hidden payloads, touching many files, triggering suspicious behaviour signatures. Dynamic analysis captures that behaviour by actually executing the app in an instrumented sandbox. The question for this project was how well models can learn from such behavioural reports, and whether a vision model can learn from an image representation of them.

The dataset: 4,280 apps through a sandbox

  • 2,280 ransomware samples from the RansomProber dataset
  • 2,000 benign apps from AndroZoo

Every APK was executed in the CuckooDroid sandbox, which runs Android apps in a controlled environment and writes a JSON report covering system calls, network activity, file operations, static APK information and matched behaviour signatures.

From a JSON report to nine features

The reports are large and nested, so the first step was feature engineering. A Python script reads each report and extracts nine behavioural features:

#FeatureSource in the report
1VirusTotal positivesvirustotal.positives
2Total signature severitysum of severity over matched signatures
3Triggered signaturesnumber of matched signatures
4Asks for dangerous permissionsbinary, from a signature
5Flagged by more than 10 antivirus enginesbinary, from a signature
6Hidden payload foundbinary, from a signature
7Dangerous permissionscount in the manifest
8Hidden payloadscount
9Flagged filescount

These nine numbers per app became a CSV for classical machine learning, and the input for the image transformation.

Behaviour to pixels

The idea behind the paper's title is to give vision models a picture of an app's behaviour. The transformation in the repository is deliberately simple:

  1. Min-max normalise the nine features of one app to the range 0–255.
  2. Assign them to colour channels by role: red for the three strongest maliciousness signals (VirusTotal positives, total severity, triggered signatures), green for the three counts (dangerous permissions, hidden payloads, flagged files) and blue for the three binary flags.
  3. Lay the values out as a 3×3 grid and upscale it to 369×369 with Lanczos resampling, which turns the nine cells into a smooth image; each model then resizes it to its own input size (224×224 for the ViT). A grayscale copy of every image is made with OpenCV.
# R: most critical, G: counts, B: binary flags
red_channel = normalized_features[[0, 1, 2]].reshape((1, 3))
green_channel = normalized_features[[6, 7, 8]].reshape((1, 3))
blue_channel = normalized_features[[3, 4, 5]].reshape((1, 3))

small_image = np.zeros((3, 3, 3), dtype=np.uint8)
small_image[:, :, 0] = np.tile(red_channel, (3, 1))
small_image[:, :, 1] = np.tile(green_channel, (3, 1))
small_image[:, :, 2] = np.tile(blue_channel, (3, 1))

Two things are worth being precise about. First, the image carries the same information as the nine numbers; it is a different representation, not more data. Second, because each app is normalised by its own minimum and maximum, the image encodes the relative profile of an app's behaviour (which signals dominate) rather than absolute magnitudes. Ransomware and benign apps turn out to have visibly different profiles, which is what the vision models pick up.

Four models

Random Forest on the CSV. scikit-learn's RandomForestClassifier on the nine features, with a stratified 80/20 train/test split and 5-fold cross-validation on the training set. Accuracy: 99.41%. That a plain tree ensemble gets this high says a lot about how separable the two classes are once the behaviour has been summarised.

CNNs on the images. A small Keras CNN (three convolution and max-pooling blocks with 16, 32 and 64 filters, then a 512-unit dense layer and a sigmoid output), trained with RMSprop on 64×64 inputs. On RGB images it reached 99.76%; on grayscale, 99.53%.

A pretrained Vision Transformer. A ViT splits an image into fixed-size patches, embeds each patch as a token, adds positional embeddings and a learnable class token, and runs the sequence through a standard transformer encoder; the class token's final state feeds the classifier. I used torchvision's ViT-B/16 with ImageNet weights as a feature extractor: the backbone is frozen and only a new two-class head is trained.

pretrained_vit_weights = torchvision.models.ViT_B_16_Weights.DEFAULT
pretrained_vit = torchvision.models.vit_b_16(weights=pretrained_vit_weights).to(device)

for parameter in pretrained_vit.parameters():   # freeze the backbone
    parameter.requires_grad = False

pretrained_vit.heads = nn.Linear(in_features=768, out_features=2).to(device)

The images go through the same preprocessing the pretrained weights expect (resize, crop to 224×224, ImageNet mean and standard deviation), and the head is trained with Adam and cross-entropy loss. A grid search over learning rate, batch size and weight decay picked lr = 0.001, batch size 64 and no weight decay. The images were split 70/20/10 into training, validation and test sets.

Results

ModelInputAccuracy
Vision Transformer (ViT-B/16)RGB images99.78%
CNNRGB images99.76%
CNNgrayscale images99.53%
Random Forest9 tabular features99.41%

The ViT's precision, recall and F1 score were each 99.76%.

How I read these numbers

All four models land between 99.4% and 99.8%. The gap between the ViT and the RGB CNN (99.78% vs 99.76%) is far too small to call one model better on a few hundred test images; a different random split could reverse it. So the claims I am comfortable making are narrower:

  • Behavioural features from dynamic analysis separate these ransomware and benign samples extremely well. That is the main finding, and the Random Forest shows it most plainly.
  • A vision model pretrained on natural photos transfers to these synthetic "behaviour images" even with its backbone frozen and only a linear head trained, which I found genuinely surprising.
  • Colour helps the CNN a little (99.76% RGB vs 99.53% grayscale), consistent with the channel assignment carrying meaning.

Near-perfect accuracy on a curated dataset is not the same as robustness in the wild, and the paper names adversarial robustness as future work. The other next steps I would take are evaluating on apps from a later time period than the training data, since malware families drift; measuring what sandbox-aware malware does to the features; and running an ablation to see which of the nine features carry the signal.

What I did

This was a research internship at NIT Kurukshetra from June to December 2024. I ran the dynamic analysis, built the Python pipeline that turns CuckooDroid's JSON into CSV features and images, trained and compared the models, and led the paper as lead and corresponding author, with my co-authors Kamaldeep and Manisha Malik. It appeared at ICDAM 2025 in London and is published in Springer's Lecture Notes in Networks and Systems.

Frequently asked questions

What accuracy did the Vision Transformer reach for Android ransomware detection?

99.78% accuracy, with 99.76% precision, recall and F1, on images generated from CuckooDroid behaviour reports of 4,280 apps (2,280 ransomware and 2,000 benign). A CNN on RGB images reached 99.76%, a CNN on grayscale images 99.53%, and a Random Forest on nine tabular features 99.41%.

How are sandbox reports turned into images?

Nine behavioural features are extracted from each CuckooDroid JSON report, normalised to 0–255 and laid out as a 3×3 RGB grid (red for the strongest maliciousness signals, green for counts, blue for binary flags), then upscaled with Lanczos resampling and resized for each model.

Which Vision Transformer was used?

torchvision's ViT-B/16 with ImageNet weights, used as a frozen feature extractor with a new two-class head trained with Adam and cross-entropy loss. A grid search chose a learning rate of 0.001, a batch size of 64 and no weight decay.

Who wrote the paper?

"From Behavior to Pixels: A Vision Transformer Approach for Android Ransomware Detection" is by Satyam Kesharwani (lead and corresponding author), Kamaldeep and Manisha Malik, published in Springer's Lecture Notes in Networks and Systems for ICDAM 2025.

Source code on GitHub Read the paper Project overview