YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
FlowResampler: Graph Diffusion Resampling over Flow-Aware Graph Representations for LLM-based Vulnerability Detection
This repository contains the implementation of FlowResampler, a graph diffusion resampling framework over flow-aware graph representations for LLM-based vulnerability detection.
Environment Setup
To set up the required environment and install all dependencies, use Conda with the provided environment.yml:
# Create the conda environment from file
conda env create -f environment.yml
# Activate the environment
conda activate <env_name>
## Results
All experiment outputs and model checkpoints are stored in the `results/` directory.
To reproduce the reported evaluation results, simply run:
```bash
python research_questions.py
The generated summaries will be saved under:
results/summary/
Training FlowResampler
To train FlowResampler from scratch, the Code Property Graphs (CPGs) and their node/edge embeddings need to be generated locally.
1. Generate Code Property Graphs
First, extract the provided Joern CLI package:
unzip utils/joern-cli.zip -d utils/
Then generate the Code Property Graphs (CPGs) for the PrimeVul test, validation, and training sets:
python -m utils.cpg_utils --file_name primevul_test.jsonl
python -m utils.cpg_utils --file_name primevul_valid.jsonl
python -m utils.cpg_utils --file_name primevul_train.jsonl
After preprocessing, the generated CPG files can be found under:
datasets/processed/primevul/data_graphs/cpg/
The .jsonl files in this directory contain the processed Code Property Graphs used by FlowResampler.
2. Initialize Node and Edge Embeddings
After generating the CPGs, initialize the node and edge embeddings by running:
python -m utils.embed_utils
The preprocessed embeddings will be stored as .pt files under:
datasets/processed/primevul/data_graphs/cpg/embeddings/
These embeddings are used as the initial graph representations during model training.
3. Train FlowResampler
Once the CPGs and embeddings have been generated, train FlowResampler using:
python trainer.py \
--experiment_name flow_resampler \
--diffusion_type bidirectional \
--num_hops 4 \
--max_num_queries 64 \
--model_name 'Qwen/Qwen2.5-Coder-3B'
The training outputs and checkpoints will be saved under:
results/flow_resampler/
Output Summary
results/.../predictions.jsonโ prediction results used for evaluation.results/summary/โ evaluation summaries generated byresearch_questions.py.results/flow_resampler/โ FlowResampler training outputs and checkpoints.datasets/processed/primevul/data_graphs/cpg/.../*.jsonlโ preprocessed CPGs.datasets/processed/primevul/data_graphs/cpg/embeddings/*.ptโ precomputed node and edge embeddings.