Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions task-submissions/haoran/1-x-1/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
/data/
/models/
__pycache__/
*.py[cod]
452 changes: 452 additions & 0 deletions task-submissions/haoran/1-x-1/ICSI_LICENSE.html

Large diffs are not rendered by default.

129 changes: 129 additions & 0 deletions task-submissions/haoran/1-x-1/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,129 @@
# Compact Conversational Memory Question Answering

Build a compact memory and a retrieval-augmented system that answers questions about meeting transcripts.

**Task:** `task-1-x-1` · **Mode:** Implementation · **Metric:** EvidenceGroundedAnswerAccuracy

## Overview

This task covers 67 meetings from three ICSI series: Bmr (29), Bro (23) and
Bed (15). The history contains 53,600 utterances and 719,788 whitespace-delimited
words, including interruptions, references to earlier discussions, corrections
and changing proposals.

The agent receives the complete history during development and builds a
finished memory, a retriever and an answerer. Evaluation withholds the original
transcripts from all three submitted programs. Each answer must be supported by the
records retrieved from the compact memory.

## What This Task Tests

- Preserving useful information from conversations within a storage budget.
- Retrieving question-relevant evidence from a locally prepared memory.
- Producing answers grounded in the retrieved records.
- Packaging retrieval and answering for separate execution environments.

## Task Setup

### Provided Assets

| Asset | Purpose |
| --- | --- |
| `data/history.jsonl` | Complete transcripts of 67 meetings |
| `data/validation/queries.jsonl` | 30 public development questions |
| `data/validation/golden_answers.jsonl` | Public reference answers and required facts |
| `data/validation/evidence.jsonl` | Transcript excerpts supporting the public answers |
| `environment/docs/` | Installed runtime and permitted API resources |

The download script restores the task's `data/` directory, mounted read-only
under `/task/data`. The 118 held-out questions concern the same history and are
disjoint from the public examples. Their answers and 1,344 supporting excerpts
are packaged under `tests/data/` and excluded from the agent environment.
Evidence provenance is checked against the supplied transcripts.

ICSI stands for **International Computer Science Institute**. The
[ICSI Meeting Corpus](https://groups.inf.ed.ac.uk/ami/icsi/download/) contains
recorded research meetings and human transcripts; see Janin et al., *The ICSI
Meeting Corpus*, ICASSP 2003. The original license notice is included.
The 148 QA examples were authored for this task and are not original ICSI
annotations. They have transcript-linked evidence but no independent human
certification. Semantic grading remains subject to model variability.

Public assets are pinned in [assets.json](assets.json) to the
[development dataset](https://huggingface.co/datasets/hrjinbb12345/search-swe-development/tree/7fe4b0bfb7699cb393bc5b11835c47ffeb13aaec).

### Fixed Components and Allowed Changes

The history, submission interfaces and resource limits are fixed. Memory
format, information selection, indexing, retrieval and answer generation are
implementation choices. The task starts without a starter implementation.

Memory construction and retrieval may use local computation and Jina embedding
and reranking APIs. Only the answerer may call the allowed OpenRouter generation
models, subject to the [resource policy](environment/docs/available_resources.md).

### Environment and Resource Limits

The CPU Python 3.12 environment provides 8 CPUs, 8 GiB memory, 8 GiB storage
and no GPU. The agent has 120 minutes; the Harbor verifier phase has 90 minutes.
Index construction has 300 seconds. Across the question set, retrieval has 150 seconds and answering 1,800 seconds. Stage budgets do not extend the overall verifier budget.
Each answer permits at most two generation API calls, 2,000 output tokens per call and
2,000 Unicode characters in the final text.

`memory.json` is readable UTF-8 JSON and may occupy at most 183,981 bytes,
5% of the 3,679,621 dialogue-text bytes. Scripts are outside this memory budget
and contain corpus-independent code.

## Submission Contract

The deliverable is `/app/memory.json` plus the executable scripts `build_index.sh`, `search.sh` and `answer.sh`. The verifier builds an index from the memory, retrieves at most ten strings per question and passes those strings with the question to the answerer. The answerer returns plain text. Original transcripts are unavailable to the submitted programs during evaluation.

The full schema, stage permissions and invocation details live in [instruction.md](instruction.md).

## Evaluation

### Search Quality

A question scores one only when both retrieved evidence and answer correctness pass. The evidence judge accepts faithful summaries or paraphrases preserving at least one question-relevant reference fact; topic overlap alone does not count. The answer judge checks the reference answer and required facts. EvidenceGroundedAnswerAccuracy is the mean of these joint outcomes over 118 held-out questions. The report's 0–100 score is 100 times the normalized reward; evidence and answer accuracy are also reported separately.

### Correctness and Resource Gates

All three entry points, output formats, execution limits and the memory budget must
pass. Invalid output or a failed submission process invalidates the submission.
Empty retrieval is an evidence miss. A malformed judge response or model
service failure invalidates the measurement.

### Integrity Checks and Final Reward

Evidence and answer judging are separate from the trajectory audit. The audit checks task compliance and can set the whole reward to zero. An incomplete audit is an infrastructure failure. Submitted programs and the audit process do not receive the judges' provider credentials. The trajectory audit uses `deepseek-flash` through pinned RewardKit 0.1.7; evidence and answer grading use independently configured judge settings.

Task identity and execution settings are recorded in `task.toml`. Final numbering and official asset migration remain maintainer steps.

## Running This Task

From the repository root, follow the [quick start](../../../docs/quickstart.md)
to prepare the runtime and Docker. The [evaluation guide](../../../docs/evaluation.md)
explains agent and verifier configuration; the [asset guide](../../../docs/assets.md)
covers downloads and checksums.

Configure `JINA_API_KEY` for optional embedding and reranking, `OPENROUTER_API_KEY` for the submitted answerer, `ANSWER_JUDGE_*` for evidence and answer grading, and `VERIFIER_OPENAI_*` for the trajectory audit. The answerer selects among the four models in the [resource policy](environment/docs/available_resources.md); judge settings are separate. Credentials must remain outside the task package.

```bash
python scripts/download_assets.py --task-path task-submissions/haoran/1-x-1
bash scripts/run_task.sh --task-path task-submissions/haoran/1-x-1 --model "YOUR_AGENT_MODEL"
```

The shared launcher defaults to Codex, also supports Pi and Claude Code, and writes results under `jobs/task-submissions/haoran/1-x-1/`. Replace `YOUR_AGENT_MODEL` with the configured model. Add `--dry-run` to inspect the launch command without starting an evaluation; it does not validate credentials, assets or API access.

## Task Files

| File or directory | What to read it for |
| --- | --- |
| [instruction.md](instruction.md) | Complete agent-facing specification and executable contract |
| [task.toml](task.toml) | Task identity, artifacts, resource limits and environment variables |
| [assets.json](assets.json) | Fixed asset paths, immutable revisions and checksums |
| [Environment guide](environment/docs/environment.md) | Installed runtime and task environment |
| [Environment configuration](environment/docker-compose.yaml) | Read-only data mounts |
| [Verifier](tests/) | Execution, output validation and scoring |
| [Resource policy](environment/docs/available_resources.md) | Permitted API calls and credential handling |
| [Corpus license](ICSI_LICENSE.html) | ICSI redistribution terms and attribution |
49 changes: 49 additions & 0 deletions task-submissions/haoran/1-x-1/assets.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
{
"schema_version": 1,
"files": [
{
"path": "data/history.jsonl",
"size_bytes": 12155329,
"sha256": "d71fbbc8f233145e155131bd01b6c104a9a5c578379cb66bfa4c8c2bba54cb52",
"source": {
"repo_id": "hrjinbb12345/search-swe-development",
"repo_type": "dataset",
"revision": "7fe4b0bfb7699cb393bc5b11835c47ffeb13aaec",
"filename": "development/task-1-x-1/history.jsonl"
}
},
{
"path": "data/validation/evidence.jsonl",
"size_bytes": 74062,
"sha256": "d7419b529598c2c392fbe346b5045789a20ecb763fb336fa7140a793d5befbf7",
"source": {
"repo_id": "hrjinbb12345/search-swe-development",
"repo_type": "dataset",
"revision": "7fe4b0bfb7699cb393bc5b11835c47ffeb13aaec",
"filename": "development/task-1-x-1/validation/evidence.jsonl"
}
},
{
"path": "data/validation/golden_answers.jsonl",
"size_bytes": 11271,
"sha256": "d6d2588018a278a86d9fa9a2e7f358b3b243a8de026d97a1338ee636575fbec6",
"source": {
"repo_id": "hrjinbb12345/search-swe-development",
"repo_type": "dataset",
"revision": "7fe4b0bfb7699cb393bc5b11835c47ffeb13aaec",
"filename": "development/task-1-x-1/validation/golden_answers.jsonl"
}
},
{
"path": "data/validation/queries.jsonl",
"size_bytes": 6950,
"sha256": "1bc181d04d065275401a5972c710912ea6c77375b10389652fc83ffff83afaa1",
"source": {
"repo_id": "hrjinbb12345/search-swe-development",
"repo_type": "dataset",
"revision": "7fe4b0bfb7699cb393bc5b11835c47ffeb13aaec",
"filename": "development/task-1-x-1/validation/queries.jsonl"
}
}
]
}
18 changes: 18 additions & 0 deletions task-submissions/haoran/1-x-1/environment/Dockerfile
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
FROM docker.io/hanhainebula/search-swe-base:cpu-py3.12-1.0.0

ARG SEARCH_SWE_CODEX_VERSION=0.147.0
ARG SEARCH_SWE_CLAUDE_CODE_VERSION=2.1.273
ARG SEARCH_SWE_PI_VERSION=0.85.1
RUN npm install --global --ignore-scripts \
--registry=https://registry.npmmirror.com \
"@openai/codex@${SEARCH_SWE_CODEX_VERSION}" \
"@earendil-works/pi-coding-agent@${SEARCH_SWE_PI_VERSION}" \
&& npm install --global \
--registry=https://registry.npmmirror.com \
"@anthropic-ai/claude-code@${SEARCH_SWE_CLAUDE_CODE_VERSION}" \
&& codex --version | grep -Fx "codex-cli ${SEARCH_SWE_CODEX_VERSION}" \
&& test "$(pi --version)" = "${SEARCH_SWE_PI_VERSION}" \
&& test "$(claude --version)" = "${SEARCH_SWE_CLAUDE_CODE_VERSION} (Claude Code)"
RUN claude --help | grep -F "(low, medium, high, xhigh, max)" >/dev/null

WORKDIR /app
21 changes: 21 additions & 0 deletions task-submissions/haoran/1-x-1/environment/docker-compose.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
services:
main:
volumes:
- type: bind
source: ../data/history.jsonl
target: /task/data/history.jsonl
read_only: true
bind:
create_host_path: false
- type: bind
source: ../data/validation
target: /task/data/validation
read_only: true
bind:
create_host_path: false
- type: bind
source: ./docs
target: /task/docs
read_only: true
bind:
create_host_path: false
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
# Available Resources

The submission may use Jina embedding and reranking APIs for memory preparation, index construction and retrieval, and the OpenRouter generation API only in its answering component.

Harbor injects `OPENROUTER_API_KEY` and `JINA_API_KEY` into the development container at runtime. Read them as ordinary environment variables, for example `$OPENROUTER_API_KEY` or `os.environ["JINA_API_KEY"]`; do not source a `.env` file inside the container. Only explicitly configured variables are supplied, not the runner's entire environment.

For development self-tests, check that the credential exists without displaying its value:

```bash
if [ -z "${OPENROUTER_API_KEY:-}" ]; then
echo "OPENROUTER_API_KEY is not set" >&2
exit 1
fi
```

```python
import os

if not os.environ.get("OPENROUTER_API_KEY"):
raise RuntimeError("OPENROUTER_API_KEY is not set")
```

Credentials are runtime resources, not build-time settings. Keep them out of code, memory, indexes, prompts and logs. During evaluation, build and search use the Jina transport and the answerer uses the OpenRouter transport documented in the task instruction. Submitted processes do not receive real provider keys.

Harbor limits submission API access to `openrouter.ai` and `api.jina.ai`. The coding agent and verifier judges use separate model access, which is not an additional submission resource. Documentation links are references, not additional permitted network destinations.

```dotenv
OPENROUTER_API_KEY=<YOUR_OPENROUTER_API_KEY>
JINA_API_KEY=<YOUR_JINA_API_KEY>
```

## Retrieval resources

Embedding and reranking models may be used through Jina during memory preparation, index construction and retrieval. Use the exact model IDs and request formats supported by Jina. These endpoints are for text embedding, candidate scoring and reranking; do not use them for generation, external search or answer lookup. OpenRouter embedding and reranking endpoints are not provided for this task. Local computation remains available without Jina calls.

Jina:

- Quickstart: `https://docs.jina.ai/get-started/quickstart`
- Embedding API documentation: `https://api.jina.ai/scalar#tag/search-foundation-models/POST/v1/embeddings`
- Reranking API documentation: `https://api.jina.ai/scalar#tag/search-foundation-models/POST/v1/rerank`
- Models: `https://jina.ai/models`

Use `JINA_API_KEY` for development calls to `https://api.jina.ai/v1/embeddings` or `https://api.jina.ai/v1/rerank`. Evaluation provides the same endpoints through the task transport. Select a model supported by the corresponding Jina endpoint; this task does not fix a single retrieval model. Text inputs are supported; fetching URLs, images or external documents is not permitted. The model running the coding session is separate from these submission resources.

## Generative LLM resources

Generative calls are allowed only through the OpenRouter API, and only from `answer.sh`, using the current question and retrieved strings. The following four model IDs are the complete allowlist:

- `qwen/qwen3.6-35b-a3b` — `https://openrouter.ai/qwen/qwen3.6-35b-a3b`
- `qwen/qwen3.5-35b-a3b` — `https://openrouter.ai/qwen/qwen3.5-35b-a3b`
- `qwen/qwen3.5-9b` — `https://openrouter.ai/qwen/qwen3.5-9b`
- `qwen/qwen3-30b-a3b-instruct-2507` — `https://openrouter.ai/qwen/qwen3-30b-a3b-instruct-2507`

The fourth model ID is the OpenRouter identifier for `Qwen3-30B-A3B-Instruct-2507`. Include your selected model ID in each request; the verifier does not impose a single answer model. Different allowed models may be selected within the task's call budget.

OpenRouter generation resources:

- Quickstart: `https://openrouter.ai/docs/quickstart`
- API documentation: `https://openrouter.ai/docs/api/reference/overview`
- OpenAI-compatible endpoint: `https://openrouter.ai/api/v1/chat/completions`

### Strict allowlist and jailbreak penalty

The four model IDs above, OpenRouter endpoint and answering-only use are mandatory restrictions, not recommendations. Do not use other providers, external search or answer services, benchmark-answer datasets, or generative APIs for memory preparation or retrieval. Only the Jina embedding and reranking endpoints above are permitted retrieval APIs. Do not reuse information across answering requests.

A prohibited API call, access to unprovided evidence or bypass of the required pipeline is a task violation. If detected by the verifier or trajectory audit, it sets the entire task score to `0`, regardless of retrieval or answer quality.
7 changes: 7 additions & 0 deletions task-submissions/haoran/1-x-1/environment/docs/environment.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# CPU Docker Environment

This image provides a Conda-managed Python 3.12 environment with CPU-only PyTorch and the common search, embedding, indexing, document-processing, media, HTTP, and service packages used by the tasks. Installed Python packages include `torch`, `torchvision`, `torchcodec`, `numpy`, `transformers`, `sentence-transformers`, `FlagEmbedding`, `deepspeed`, `faiss-cpu`, `bm25s`, `rank-bm25`, `pyserini`, `hnswlib`, `qdrant-client`, `docling`, `marker-pdf`, `pdf2image`, `pypdfium2`, `CairoSVG`, `av`, `imageio`, `imageio-ffmpeg`, `tiktoken`, `fastapi`, `uvicorn`, `python-multipart`, `requests`, `aiohttp`, `openai`, and `pydantic-settings`.

The task Python interpreter and its installed packages are available at `/opt/conda/bin/python`. Use `/opt/conda/bin/python` and `/opt/conda/bin/pip` when invoking Python or installing packages.

The image also includes JDK 21, Node.js 24.16.0 with npm 11.13.0, FFmpeg, Poppler utilities, Cairo, Git, curl, `jq`, `build-essential`, `ca-certificates`, `libffi`, `libgomp`, `netbase`, `netcat`, `procps`, `tzdata`, `unzip`, and the related system runtime libraries.
Loading
Loading