input_file 로 PDF 를 직접 보내면 별도 chunking 이나 OCR pipeline 없이 내용을 읽힐 수 있어. 여기에 response_format 의 Pydantic 모델을 결합하면 짧은 코드로 invoice 나 report 추출기를 만들 수 있어.
여러 전처리 단계를 한 호출로 줄여
예전에는 pdfplumber 로 text 를 꺼내고, chunking 과 embedding, vector search, prompt 조립, JSON parsing 을 이어 붙이곤 했어. 문서가 한 번에 처리 가능한 크기라면 input_file 과 response_format 으로 이 단계를 한 API 호출에 묶을 수 있어.
아주 긴 문서는 여전히 나눠야 해
input_file 은 PDF 전체를 한 번에 처리하므로 100 페이지가 넘는 긴 문서는 context window 한도에 걸릴 수 있어. 이런 경우에는 chunking 이 필요해. invoice, report, contract, 한 장짜리 form 에는 직접 입력 방식이 잘 맞아.
Pydantic 모델이 출력 구조를 정해
class Invoice(BaseModel): vendor: str; invoice_number: str; total: float; currency: str; line_items: list[LineItem] 를 만들고 response_format=Invoice 로 호출하면 그 구조에 맞는 결과를 받아 방어적인 parsing 을 줄일 수 있어.