aqora / PQID: Parallel Quantum Instruction Dataset
import pandas as pd
df = pd.read_parquet("aqora://aqora/pqid/v1.0.2")
repo_owner/repo_name.page_size to embed more in the static export.OPENQASM 3.0;
include "stdgates.inc";
bit[3] c;
qubit[3] q;
s q[2];
z q[0];
h q[0];
train · generated by gpt-5.4 (base_seed_quality_aware) · hallucination type nonepqid. Filtering on seed_role, validation_status or
repo_owner is fastest, because the data is sorted by those columns.SELECT instruction, response, openqasm3_code, repo_owner, repo_name, repo_license, github_anchor
FROM pqid
WHERE seed_role = 'gold_generation'
LIMIT 20page_size to embed more in the static export.| Column | Description |
|---|---|
split | train, validation or test, as released |
instruction / response | Task and reference response, generated by an LLM from the source snippet |
openqasm3_code | OpenQASM 3 export of the circuit, where the source executed and export succeeded |
seed_role, seed_expected_response_mode | Task type: generation, mutation robustness, repair/explanation or diagnosis |
prompt_type, paraphrase_source, generation_model | Seed vs. paraphrase lineage and the generating model |
validation_status, hallucination_type, validation_error_type | Outcome of executing the source snippet in Qiskit |
num_qubits, circuit_depth, gate_count, gate_types, … | Circuit structure (only for executed circuits; gate_types is a JSON object) |
transpiled_*, transpilation_* | Transpilation statistics |
repo_owner, repo_name, original_url, github_anchor, start_line, end_line | Source provenance |
repo_license, license_category, distribution_rights_status, license_previous_state | Licence evidence and release governance |
content_hash | Unique row id |
license_category = 'permissive'). A small number of these
come from LGPL, MPL or EUPL repositories, so check repo_license and keep upstream
notices if you redistribute the code. The authors' attribution manifest is in the
Zenodo archive.
Changes in this copy: the train, validation and test JSONL files are merged into one
Parquet table with a split column; input and output are renamed to instruction
and response; the nested metadata object is flattened into columns
(gate_types and license_previous_state stay as JSON strings); rows are sorted
for faster filtering. Row content is unchanged.