ParEval

This repo contains the Parallel Code Evaluation (ParEval) Benchmark for evaluating the ability of Large Language Models to write parallel code. See the ParEval Leaderboard for up-to-date results on different LLMs. We have extended this to include testing for HPX and Legion generation, and in addition translation from HPX -> Legion

Overview

The organization of the repo is as follows.

prompts/ -- the prompts in ParEval alongside some utility scripts
generate/ -- scripts for generating LLM outputs
drivers/ -- scripts to evaluate LLM outputs
analysis/ -- scripts to analyze driver results and compute metrics
- @k/ -- summary csvs of HPX generation performance
- visuals/ -- per-prompt category net runtime graphs
- visuals_specific/ -- individual prompt runtime graphs
tpl/ -- git submodule dependencies
prompts/ -- all prompt file jsons that are to be used as inputs for the generation phase
run_driver -- miscellaneous driver running scripts
srun_generate/ -- miscellaneous code generation scripts

Each subdirectory has further documentation on its contents. The general workflow is to use generate/generate.py to generate LLM outputs, run drivers/run-all.py to evaluate outputs, and analysis/metrics.py to post-process the results summaries.

Setup and Installation

A couple core systems software are assumed to be installed: Python >=3.7, a C++ compiler that supports C++17 and OpenMP, Make, CMake, and an MPI implementation. If you are testing the CUDA and HIP prompts, then you will need access to NVIDIA and AMD GPUs alongside their respective software stacks.

First, clone the repo.

git clone https://github.com/SanjanaYasna/ParEval_amt.git

Next, you need to build Kokkos (if you want to include it in testing). Kokkos source code was installed under directory ParEval_amt/tpl/kokkos/kokkos, set to version 4.5.01

cd amt/tpl/kokkos/kokkos
git checkout {stable version of choice}
module load gcc/9.4.0
module load mpich/4.2.1
#configure again at your own preferences, with build set to builddir below 
#my config...
cmake -B builddir \
    -DCMAKE_CXX_COMPILER=g++ \
    -DCMAKE_BUILD_TYPE=Release \
    -DKokkos_ENABLE_OPENMP=ON \
    -DKokkos_ENABLE_THREADS=ON \
    -DKokkos_ARCH_NATIVE=ON \
    -DKokkos_ENABLE_DEPRECATED_CODE_4=OFF
cmake --build builddir
#set install prefix for kokkos to build under tpl/kokkos, as that's where pareval make files check
cmake --install builddir /work/pi_mrobson_smith_edu/ParEval_amt/tpl/kokkos/build

You will need to be able to use HPX for this project. There are two versions of HPX being tested: 1.5.1, and 1.10.0 There are setup scripts on the Unity cluster to get the respective version of HPX running:

#HPX 1.5.1 
source /work/pi_mrobson_smith_edu/.hpx_tcmalloc_1.5.1
#OR 
#HPX 1.10.0 
source /work/pi_mrobson_smith_edu/.hpx_1_10_0

Finally, you need to install the Python dependencies. requirements_AMT.txt has the set of dependencies. Use UV for the easiest time installing these.

#get uv in whatever environment you have
pip install uv
#if you're on unity cluster, there is a uv environment you can activate
source /work/pi_mrobson_smith_edu/pareval/.venv/bin/activate

#otherwise, take from the .txt environment file and make a uv environment from these packages 
uv add -r requirements_AMT.txt

Citing ParEval original repo contents

@misc{nichols2024large,
      title={Can Large Language Models Write Parallel Code?}, 
      author={Daniel Nichols and Joshua H. Davis and Zhaojun Xie and 
              Arjun Rajaram and Abhinav Bhatele},
      year={2024},
      publisher = {Association for Computing Machinery},
      address = {New York, NY, USA},
      booktitle = {Proceedings of the 33rd International Symposium on High-Performance Parallel and Distributed Computing},
      series = {HPDC '24}
}

License

ParEval is distributed under the terms of the MIT license.

Name		Name	Last commit message	Last commit date
Latest commit History 67 Commits
.github/workflows		.github/workflows
analysis		analysis
bin		bin
drivers		drivers
generate		generate
high_reasoning_scripts		high_reasoning_scripts
prompts		prompts
run_driver		run_driver
srun_generate		srun_generate
tpl		tpl
.gitignore		.gitignore
.gitmodules		.gitmodules
CITATION.cff		CITATION.cff
LICENSE		LICENSE
README.md		README.md
generate_gemini.sh		generate_gemini.sh
generate_locking_contention.sh		generate_locking_contention.sh
generate_locking_contention_gpt.sh		generate_locking_contention_gpt.sh
hpx_driver_test_la.sh		hpx_driver_test_la.sh
hpx_driver_test_magicoder.sh		hpx_driver_test_magicoder.sh
hpx_driver_test_reduce.sh		hpx_driver_test_reduce.sh
hpx_singular.sh		hpx_singular.sh
hpx_tcmalloc_gr_red_la.sh		hpx_tcmalloc_gr_red_la.sh
og_adj.sh		og_adj.sh
one_histogram.sh		one_histogram.sh
requirements_AMT.txt		requirements_AMT.txt
run_fft.sh		run_fft.sh
run_geometry.sh		run_geometry.sh
run_histogram.sh		run_histogram.sh
run_scan.sh		run_scan.sh
run_search.sh		run_search.sh
run_sort.sh		run_sort.sh
run_stencil.sh		run_stencil.sh
run_transform.sh		run_transform.sh
run_transform_adj.sh		run_transform_adj.sh
sani		sani
sanity.json		sanity.json
sanity.sh		sanity.sh
scan_adjusted.sh		scan_adjusted.sh
test_hpx_cpu.sh		test_hpx_cpu.sh
test_hpx_cpu_batch.sh		test_hpx_cpu_batch.sh
test_hpx_gpu.sh		test_hpx_gpu.sh

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

ParEval

Overview

Setup and Installation

Citing ParEval original repo contents

License

About

Uh oh!

Releases

Packages

Languages

License

mpr-lab/ParEval_AMT

Folders and files

Latest commit

History

Repository files navigation

ParEval

Overview

Setup and Installation

Citing ParEval original repo contents

License

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages