[DL 2026] 1. Project Settings
Published:
In this post, Project Settings for Introduction to Deeplearning class is introduced.
Code Structure
/dl2026/
dataset/ # 제공 public dataset, 건드리지 않음
scripts/
setup.sh # 테스트용 wrapper
evaluate.sh # 테스트용 wrapper
skeleton/ # 시작 템플릿
src/
setup.sh
evaluate.py
pyproject.toml
uv.lock
/workspace/project/
src/
__init__.py
solver.py # 핵심 예측 코드
setup.sh # 평가 환경 준비
pyproject.toml # 의존성
uv.lock # uv lock 파일
evaluate.py # 로컬 테스트용, 제출에는 의존 X
artifacts/ # 선택: 모델 weight, spec 요약, 캐시 등
setup.sh: 20분간의 setup phase에서 dependay 설치, HuggingFace 모델 캐싱, spec preprocessing, retrieval index 생성, rule cache 생성 등의 작업을 하도록 직접 script 작성.pyproject.toml: 프로젝트에 필요한 라이브러리를 dependency에 정의uv.lock: 정확히 어떤 버전을 설치할지 정의.uv lock커맨드를 수행하면uv.lock생성uv synd: virtual environment(.venv) 생성하고uv.lock을 기준으로 패키지 설치uv run python evaluate.py하면 현재 프로젝트 루트에서 만들어진 가상환경으로 python 실행현재 프로젝트(current project)는 보통
pyproject.toml을 기준으로 uv가 찾은 프로젝트 루트 디렉터리를 의미한다. 즉 uv는 현재 작업 디렉터리(pwd)부터 시작해서현재 폴더 확인
없으면 부모 폴더로 올라감
pyproject.toml찾을 때까지 반복
해서 가장 가까운
pyproject.toml이 있는 디렉터리를 프로젝트 루트로 간주한다.uv run python -c "import sys; print(sys.executable)"커맨드를 통해 현재 디렉터리에서 uv가 프로젝트 디렉터리로 무엇을 보고 있는지 확인 가능하다.
evaluate.py:#!/usr/bin/env python3 import json import os from pathlib import Path ## __file__ : 현재 실행 중인 파이썬 파일의 경로 ## ROOT : 현재 파일이 위치한 디렉터리 ROOT = Path(__file__).resolve().parent # evaluate.py가 있는 폴더를 기준으로 src/solver.py 안의 Solver 클래스를 가져옴 from src.solver import Solver # 환경변수 DATASET_DIR 이 있으면 그 값을 쓰고, 없으면 /worksapce/dataset 사용 _DATASET_ROOT = Path(os.environ.get("DATASET_DIR", "/workspace/dataset")) DATASET_DIR = _DATASET_ROOT / "testcases" # 환경변수 LABEL_PATH 가 있으면 그 값을 쓰고, 없으면 _DATASET_ROOT/"label1.jsonl" 사용 LABEL_PATH = Path(os.environ.get("LABEL_PATH", _DATASET_ROOT / "label.jsonl")) # 파일 이름에서 testcase 번호 추출 ## tc12.json -> 12 def case_number(path): return int(path.stem.removeprefix("tc").split("_")[0]) # JSON 파일을 읽어서 파이썬 객체(dict/list 등)로 변환 ## 파일 안 JSON 문자열을 읽어서 파이썬 객체로 변환하여 리턴 def load_json(path): with path.open() as f: return json.load(f) # dataset_dir 안에 있는 testcase JSON 파일들을 전부 읽어서 정렬 후 리스트 형태로 반환 ## dataset_dir.glob("tc*.json") : tc로 시작하고 .json으로 끝나는 파일 모두 찾아 iterator 반환 def load_dataset(dataset_dir): """Load all testcases as [{"id": str, "steps": list}, ...] sorted by case number.""" return [ {"id": path.name, "steps": load_json(path)} for path in sorted(dataset_dir.glob("tc*.json"), key=case_number) ] # "파일명 -> label" 형태의 dict를 만드는 함수 ## 입력 (jsonl 파일) ### {"filename":"tc1.json","label":1} ### {"filename":"tc2.json","label":0} ## 출력 ### { ### "tc1.json": 1, ### "tc2.json": 0 ### } def load_labels(path): """JSONL: one {"filename": ..., "label": ...} record per line.""" labels = {} with path.open() as f: for line in f: line = line.strip() if not line: continue record = json.loads(line) labels[record["filename"]] = record["label"] return labels def main(): dataset = load_dataset(DATASET_DIR) labels = load_labels(LABEL_PATH) solver = Solver() predictions = solver.predict(dataset) correct = 0 total = 0 with open("predictions.jsonl", "w") as pred_file: for item in dataset: case_id = item["id"] prediction = predictions.get(case_id, "fail") answer = labels[case_id].strip().lower() correct += int(prediction == answer) total += 1 pred_file.write(json.dumps({"id": case_id, "prediction": prediction}) + "\n") score = 100.0 * correct / total if total else 0.0 with open("scores.json", "w") as score_file: json.dump({"score": score}, score_file) score_file.write("\n") print(f"score={score:.2f}") if __name__ == "__main__": main()파이썬은
from src.solver import Solver를 볼 때src라는 패키지를 어디서 찾지? 라는 질문에 대해sys.path안에 있는 경로들을 기준으로 탐색한다.python evaluate.py으로 파이썬을 실행하면sys.path[0]에 evaluate.py가 있는 디렉터리가 자동 추가된다.패키지로 사용할 디렉터리 아래에
__init__.py를 추가하면 해당 디렉터리를 패키지로 인식하여 import 가 가능해진다. 즉,sys.path안에 디렉터리들 중 이것이 패키지인가를 우선적으로 확인하는데 이때__init__.py의 존재 여부가 필요하다. (python 3.3이상에서는 필요없지만 여전히 관례적으로 사용한다.)# dataset = load_dataset(DATASET_DIR) dataset = [ { "id": "tc1.json", "steps": [ { "index": 1, "input": {...}, "output": {...} } ] }, { "id": "tc2.json", "steps": [...] }, ... ] # labels = load_labels(LABEL_PATH) labels = { "tc1.json": "pass", "tc2.json": "fail", "tc3.json": "pass" } # solver = Solver() # predictions = solver.predict(dataset) { "tc1.json": "PASS", "tc2.json": "FAIL" }
❓extract_states.py 에 main에 base="./dataset" 에서, dataset 디렉터리가 항상, src 디렉터리와 같은 계층에 있다고 가정하여도 되나?
To Do
- Branch 관계 파악
- fine-tuning branch main에 반영된건지 : 반영 안됨.
- 현재 리더보드 점수 파악
- simulator 코드 파악
- 준서 generator.py (simulator) 용도를 현재는 데이터셋을 만들어서, rule-based solver 결과와 비교해서, rule-based 구현의 타당도를 평가하는 데에만 쓰였는지, 이를 학습 데이터로 쓰면 어떨지?
- 그렇게 하고 있음.
- 현재 시뮬레이터로 2만개 데이터중 룰 베이스로 안걸러지는 약 5천개 저장. -> rag 모델로 64.5점을 기반으로 계산한 결과 근접하면 데이터 성능 좋은것.
Rag













evaluate.py
↓
solver.py
↓
prompt_parser.py
↓
extract_states.py
↓
query_builder.py
↓
retriever.py
├─ BM25 검색
└─ Embedding 검색
↓
Hybrid 점수 결합
↓
관련 문서 Top-k 추출
↓
solver.py
↓
최종 PASS / FAIL
submit







generate.py


Leave a Comment