
Paper · EMNLP 2025 · NLLP Workshop
Translating Tax Law to Code with LLMs
A Benchmark and Evaluation Framework
- Built a dataset for legal code generation by collecting Catala code published online.
- Defined evaluation metrics for Catala, such as CodeBLEU and tree-edit distance.
- Evaluated LLM families such as Llama and Phi with these custom metrics.




