Hi, thanks for releasing PailGen.
The README says the hybrid retriever's output is retrieved_results_bigvul_cvefixes_top50.json (top 50 pairs per sample), but the dataset shared on Google Drive only contains retrieved_results_bigvul_cvefixes_top10.json and retrieved_results_d2a_top10.json (10 pairs per sample), and the repository itself ships no retrieval results.
Since generate_patterns.py takes list(templates.items())[:10] and process_prompt_data.py needs 10 fix patterns per sample, the top-10 version isn't sufficient and the pipeline fails with an IndexError.
Could you please share retrieved_results_bigvul_cvefixes_top50.json and retrieved_results_d2a_top50.json?
Thanks a lot for your time.
Hi, thanks for releasing PailGen.
The README says the hybrid retriever's output is
retrieved_results_bigvul_cvefixes_top50.json(top 50 pairs per sample), but the dataset shared on Google Drive only containsretrieved_results_bigvul_cvefixes_top10.jsonandretrieved_results_d2a_top10.json(10 pairs per sample), and the repository itself ships no retrieval results.Since
generate_patterns.pytakeslist(templates.items())[:10]andprocess_prompt_data.pyneeds 10 fix patterns per sample, the top-10 version isn't sufficient and the pipeline fails with anIndexError.Could you please share
retrieved_results_bigvul_cvefixes_top50.jsonandretrieved_results_d2a_top50.json?Thanks a lot for your time.