# AgriGround Annotations AgriGround contains 11,421,148 Stage 4 instruction-tuning records across 503,919 source images and 14 tasks. | Split | Source images | Annotation records | |---|---:|---:| | Train | 401,234 | 9,095,320 | | Test | 102,685 | 2,325,828 | | **Total** | **503,919** | **11,421,148** | ## Layout The `train/` and `test/` directories are grouped by source family and dataset. Each leaf directory contains a gzip-compressed `stage4.jsonl.gz` file. `manifest.json` provides row counts, task counts, compressed sizes, and SHA-256 checksums for every file. `validation.json` records the full-release integrity and schema validation result. ## Record Format Each line is one JSON object with these common fields: - `image_path`: path of the image in its source dataset - `task_type`: machine-readable task identifier - `question_type`: human-readable task identifier - `question`: model instruction or question - `answer`: target response Task-dependent fields include COCO-style RLE masks, normalized `[x, y, width, height]` bounding boxes, object counts, grounded phrases, regions, and multi-turn conversations. The 14 task definitions and exact distribution are listed in [`docs/TASKS.md`](../docs/TASKS.md). ## Loading ```python import gzip import json with gzip.open("annotations/train/classification/banana_leaf_disease_classification/stage4.jsonl.gz", "rt", encoding="utf-8") as stream: for line in stream: record = json.loads(line) print(record["task_type"], record["question"], record["answer"]) break ``` Source images are not included. Resolve each `image_path` against the corresponding source dataset and follow that dataset's original license and usage terms.