MS capstone · RIT · two semesters · team of three
Security-Aware Speculative Decoding
Code written so far
promptaccepted tokens
Draft and verify
Standard speculative decoding
Small model drafts
Llama 3.2 1B5 tokens
Cheap guesses at what comes next
Large model checks
Llama 3.2 3Bone pass
Scores all five drafted tokens at once
Out: how much the large model agrees with each drafted token
Security check
The part this project adds
Wait for a full statement
statement detector
Passes it on with 5 lines of context
Score it
CodeBERTfine-tuned
How likely is this statement to be secure?
Out: a modifier from 0.05 (penalise) to 2.0 (reward)
Accept or reject each drafted token
agreement × security modifier
Insecure-looking code makes the draft harder to accept, so the large model rewrites it. Then the round repeats.
The problem
Language models that write code also reproduce the insecure habits in the code they learned from: SQL built by gluing strings together, shell commands run with shell=True, pickle used on untrusted input. The usual fixes are to retrain the model, which is expensive, or to ask for secure code in the prompt, which the model is free to ignore. Our capstone asks a different question: can a model be steered toward safer code while it is generating, without changing its weights at all?
How we approached it
- 01
Explore. Over two semesters we worked through several designs before settling on this one, including blending the two models' logits and fine-tuning the drafter on secure code with LoRA. Fine-tuning cut vulnerabilities by 60% in our earlier experiments but ties the method to one model, so we looked for an approach that leaves every language model untouched.
- 02
Idea. Speculative decoding already has a small model draft tokens and a large model accept or reject them. We multiply that acceptance probability by a security score, so drafts that follow insecure-looking code are harder to accept and the large model rewrites them.
- 03
Dataset. Scanned 457,461 Python functions from CodeSearchNet with 19 patterns covering weaknesses such as SQL injection, command injection, unsafe deserialization, hardcoded credentials, weak cryptography and path traversal, along with their safe counterparts. Each match is kept with three lines of context either side, giving 11,082 fragments, half secure and half insecure, split 70/15/15.
- 04
Classifier. Fine-tuned CodeBERT (124M parameters) with a small classification head for 5 epochs at a learning rate of 2e-5 with warmup, then calibrated its confidence with temperature scaling so the score can be used directly as a multiplier.
- 05
Decoder. Implemented speculative decoding from scratch in PyTorch: Llama 3.2 1B drafts 5 tokens, Llama 3.2 3B verifies all of them in one forward pass, each is accepted or rejected on the probability ratio, and a correction token is drawn after a rejection.
- 06
Steering. A statement detector watches the token stream and fires when a complete Python statement has been written. The classifier scores it with five lines of context, and the score becomes a modifier between 0.05 and 2.0 that applies to every draft token until the next statement.
- 07
Evaluation. Ran the 121 prompts of the SecurityEval benchmark with and without steering and scanned every output with Bandit, a third-party static analyser that knows nothing about our classifier or its training data.
The hard part
Getting the classifier to behave during real generation took two attempts. The first version was trained on Big-Vul, a C and C++ vulnerability dataset, and scored every piece of Python as insecure, so almost every drafted token was rejected. Retraining on Python fixed that but exposed a subtler problem: the new classifier scored 99.4% on its test set, yet rated ordinary code, an import or a loop, as insecure, with scores between 0.01 and 0.15, because it had only ever seen security-relevant lines. With the natural cut-off of 0.5 the modifier penalised nearly everything the model wrote. Moving the threshold to 0.2 and calibrating the classifier's confidence made it react only to statements that look genuinely risky.
Results
75.8% fewer flagged completions
On the SecurityEval benchmark of 121 security-sensitive prompts, the share of completions flagged by Bandit fell from 3.3% to 0.8% with steering switched on.
Completions flagged by Bandit on SecurityEval Without steering3.3%With steering0.8%No retraining, and any model pair
Both language models stay frozen; the only thing trained is a small classifier. The method needs a drafter and a target that share a tokenizer, so the same code runs with another model family by changing two lines.
Selective, with a small cost in speed
The classifier scored about 8 statements per prompt and 57.5% of drafted tokens were still accepted, so it intervenes only where code looks risky. Generation ran at 34.3 tokens a second against 37.6 for the large model on its own, an 8.8% overhead.
Where it goes next
The clearest next step is training the classifier on vulnerabilities confirmed in real CVE fixes, in place of pattern-matched labels. Beyond that: scoring individual tokens instead of whole statements, stronger penalties for more severe weakness types, and evaluation on benchmarks with higher baseline vulnerability rates, such as CyberSecEval, using additional analysers like Semgrep and CodeQL.
- PyTorch
- Hugging Face Transformers
- Llama 3.2
- CodeBERT
- Speculative decoding
- Bandit
- CUDA