Nuitka decompiler eval

dev

holdout

runpassfailerrortotalwall seconds
dev/cruxeval717071788473functions
dev/humaneval_plus16008168132functions
dev/thealgorithms1832047523071866functions
dev total2709055432632471
holdout/mbpp_plus166037203126functions
holdout total166037203126

pass: the source with the decompiled function spliced in passed the checks of the task.
fail: that source did not pass them (a failing check, a check that timed out, or no def of the output to splice in).
error: the decompiler produced no Python (it raised, or its worker was killed or timed out).
dev, of the errors: 0 WorkerKilled, 0 FunctionTimeout.
holdout, of the errors: 0 WorkerKilled, 0 FunctionTimeout.
wall seconds: from the start to the end of the run, with the workers in parallel; a total is the sum of the runs of that corpus.

passfailerror

Readability

Python

The passed functions, as the original def and as the decompiled output (ours). The output of a function that did not pass means something else than the original, so it is not averaged. avg: per function.

runpassavg lines (original -> ours)avg cognitive complexity (original -> ours)
dev/cruxeval7175.3 -> 6.21.8 -> 1.9
dev/humaneval_plus16016.9 -> 18.93.1 -> 3.9
dev/thealgorithms183218.8 -> 20.63.1 -> 4.1
dev total270915.1 -> 16.72.7 -> 3.5
holdout/mbpp_plus1663.5 -> 4.91.3 -> 1.5
holdout total1663.5 -> 4.91.3 -> 1.5

Pseudo-C

Every function that can be compared with angr, passed or not. avg: per function.

runfunctionstotalboilerplate share of angrsum boilerplate (angr -> ours)sum python ops (angr -> ours)avg lines (angr -> ours)avg cognitive complexity (angr -> ours)
dev/cruxeval78778885.4%45705 -> 3467843 -> 82269.5 -> 28.9179.8 -> 3.9
dev/humaneval_plus16816883.3%10973 -> 52194 -> 12306.6 -> 32.6235.4 -> 5.3
dev/thealgorithms2297230780.3%185680 -> 262045496 -> 833420.9 -> 46.7447.3 -> 9.0
dev total3252326381.4%242358 -> 297155533 -> 927378.3 -> 41.7371.6 -> 7.6
holdout/mbpp_plus20320386.2%9289 -> 21484 -> 32202.2 -> 22.0145.0 -> 2.3
holdout total20320386.2%9289 -> 21484 -> 32202.2 -> 22.0145.0 -> 2.3

functions: those of the total with a pseudo-C of the last pass and angr values, whatever their status; the rest has none (its worker died, the binary failed to load, or the pseudo-C could not be rendered).
angr: the C of angr's own decompiler, without the passes of this project. ours: the pseudo-C after the last pass.
boilerplate: work CPython does on its own when the recovered source runs (reference counting, frames, exception propagation); every store, and the calls of a fixed list of runtime functions checked against Nuitka's code generation.
python ops: calls that carry a Python operation but no pass lifted to a marker.
boilerplate share of angr: boilerplate / (boilerplate + python ops) in angr's output.

Distribution

One dot per passed function, larger where several functions share the values; a dot on the line "same" is as large as the original, one above "2x" is more than twice as large. The axes are square-root scaled. A dot opens the function it stands for, or the functions with those values where several share them.

lines

dev
passed functions: original vs ours
median 10 -> 12
holdout
passed functions: original vs ours
median 2 -> 3

cognitive complexity

dev
passed functions: original vs ours
median 1 -> 2
holdout
passed functions: original vs ours
median 0 -> 0

Datasets

The original Python sources are the work of the authors of each dataset and are subject to the licenses of the dataset.

datasetsourcelicenses
cruxevalhttps://github.com/facebookresearch/cruxeval/tree/10d71b4ca99f2b6624bef9533ade5a0594ab7716MIT
humaneval_plushttps://github.com/evalplus/humanevalplus_release/tree/68cd26d53a0dec69f85eafe1f82a2a74155a2bd6Apache-2.0, MIT
mbpp_plushttps://github.com/evalplus/mbppplus_release/tree/64fc4195b858a17cdfdb3324f0baf37939144e14Apache-2.0, CC-BY-4.0
thealgorithmshttps://github.com/TheAlgorithms/Python/tree/f8084d92b2df4faecad57e0cd8c5ab2adbbf8b19MIT