YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

FlowResampler: Graph Diffusion Resampling over Flow-Aware Graph Representations for LLM-based Vulnerability Detection

This repository contains the implementation of FlowResampler, a graph diffusion resampling framework over flow-aware graph representations for LLM-based vulnerability detection.

Environment Setup

To set up the required environment and install all dependencies, use Conda with the provided environment.yml:

# Create the conda environment from file
conda env create -f environment.yml

# Activate the environment
conda activate <env_name>

## Results

All experiment outputs and model checkpoints are stored in the `results/` directory.

To reproduce the reported evaluation results, simply run:

```bash
python research_questions.py

The generated summaries will be saved under:

results/summary/

Training FlowResampler

To train FlowResampler from scratch, the Code Property Graphs (CPGs) and their node/edge embeddings need to be generated locally.

1. Generate Code Property Graphs

First, extract the provided Joern CLI package:

unzip utils/joern-cli.zip -d utils/

Then generate the Code Property Graphs (CPGs) for the PrimeVul test, validation, and training sets:

python -m utils.cpg_utils --file_name primevul_test.jsonl
python -m utils.cpg_utils --file_name primevul_valid.jsonl
python -m utils.cpg_utils --file_name primevul_train.jsonl

After preprocessing, the generated CPG files can be found under:

datasets/processed/primevul/data_graphs/cpg/

The .jsonl files in this directory contain the processed Code Property Graphs used by FlowResampler.

2. Initialize Node and Edge Embeddings

After generating the CPGs, initialize the node and edge embeddings by running:

python -m utils.embed_utils

The preprocessed embeddings will be stored as .pt files under:

datasets/processed/primevul/data_graphs/cpg/embeddings/

These embeddings are used as the initial graph representations during model training.

3. Train FlowResampler

Once the CPGs and embeddings have been generated, train FlowResampler using:

python trainer.py \
    --experiment_name flow_resampler \
    --diffusion_type bidirectional \
    --num_hops 4 \
    --max_num_queries 64 \
    --model_name 'Qwen/Qwen2.5-Coder-3B'

The training outputs and checkpoints will be saved under:

results/flow_resampler/

Output Summary

  • results/.../predictions.json โ€” prediction results used for evaluation.
  • results/summary/ โ€” evaluation summaries generated by research_questions.py.
  • results/flow_resampler/ โ€” FlowResampler training outputs and checkpoints.
  • datasets/processed/primevul/data_graphs/cpg/.../*.jsonl โ€” preprocessed CPGs.
  • datasets/processed/primevul/data_graphs/cpg/embeddings/*.pt โ€” precomputed node and edge embeddings.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support