-
Notifications
You must be signed in to change notification settings - Fork 1
[Feat] 벡터 검색 인프라 구축 - OpenSQL+pgvector 및 bge-m3 임베딩 서버 #38
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
5 commits
Select commit
Hold shift + click to select a range
3e500e4
chore: OpenSQL 위에 pgvector 0.8.0 컴파일 설치 및 docgrid DB 자동 초기화 구성
kangcheolung 5bd192d
feat: bge-m3 임베딩 서버 추가 (FastAPI + BAAI/bge-m3 1024차원)
kangcheolung 6efcfb2
docs: 벡터 검색 인프라 및 임베딩 서버 구축 문서 작성 (#35)
kangcheolung 676b376
Merge remote-tracking branch 'origin/develop' into feature/35
kangcheolung f4a6cd5
fix: CodeRabbit 리뷰 반영 - healthcheck start_period 증가, pg_hba.conf scra…
kangcheolung File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,16 @@ | ||
| FROM tmaxopensql/postgres:14.6 | ||
|
|
||
| USER root | ||
|
|
||
| RUN curl -L https://github.com/pgvector/pgvector/archive/refs/tags/v0.8.0.tar.gz \ | ||
| -o /tmp/pgvector.tar.gz \ | ||
| && tar -xzf /tmp/pgvector.tar.gz -C /tmp \ | ||
| && cd /tmp/pgvector-0.8.0 \ | ||
| && make PG_CONFIG=/usr/pgsql-14/bin/pg_config \ | ||
| && make install PG_CONFIG=/usr/pgsql-14/bin/pg_config \ | ||
| && rm -rf /tmp/pgvector.tar.gz /tmp/pgvector-0.8.0 | ||
|
|
||
| COPY init-and-start.sh /usr/local/bin/init-and-start.sh | ||
| RUN chmod +x /usr/local/bin/init-and-start.sh | ||
|
|
||
| CMD ["/usr/local/bin/init-and-start.sh"] |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,24 @@ | ||
| #!/bin/bash | ||
| set -e | ||
|
|
||
| PGDATA="${PGDATA:-/var/lib/pgsql/14/data}" | ||
|
|
||
| # Temporarily start postgres (local socket only) for initialization | ||
| pg_ctl start -D "$PGDATA" -l /tmp/pg_init.log -o "-h ''" -w | ||
|
|
||
| # Create docgrid database if not exists | ||
| psql -d postgres -tc "SELECT 1 FROM pg_database WHERE datname = 'docgrid'" | grep -q 1 \ | ||
| || psql -d postgres -c "CREATE DATABASE docgrid OWNER docgrid;" | ||
|
|
||
| # Enable pgvector in docgrid database | ||
| psql -d docgrid -c "CREATE EXTENSION IF NOT EXISTS vector;" | ||
|
|
||
| # Allow TCP connections from any host (needed for host-machine Spring Boot) | ||
| grep -qxF "host all all 0.0.0.0/0 scram-sha-256" "$PGDATA/pg_hba.conf" \ | ||
| || echo "host all all 0.0.0.0/0 scram-sha-256" >> "$PGDATA/pg_hba.conf" | ||
|
|
||
| # Stop temp postgres cleanly before handing off | ||
| pg_ctl stop -D "$PGDATA" -m fast -w | ||
|
|
||
| # Start postgres in foreground (replaces this process) | ||
| exec postgres |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -1,10 +1,6 @@ | ||
| pg_owner: app | ||
| pg_group: app | ||
| pg_superuser: app | ||
| pg_superuser_password: local_password | ||
| pg_owner: docgrid | ||
| pg_group: docgrid | ||
| pg_superuser: docgrid | ||
| pg_superuser_password: docgrid1234 | ||
| pg_database: postgres | ||
| pg_databases: | ||
| - name: app | ||
| owner: app | ||
| encoding: UTF-8 | ||
| use_system_user: false |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,181 @@ | ||
| # 벡터 검색 인프라 구축 (#35) | ||
|
|
||
| ## 1. OpenSQL + pgvector 구성 | ||
|
|
||
| ### 배경 | ||
|
|
||
| 벡터 검색 기능 구현에 pgvector 확장이 필요하다. | ||
| 공식 `pgvector/pgvector` Docker 이미지가 있지만, 대회 요구사항이 **OpenSQL(tmaxopensql) 기반**이라 교체할 수 없었다. | ||
| `tmaxopensql/postgres:14.6` 이미지에는 pgvector가 포함되어 있지 않아서 직접 컴파일해서 설치하는 방식을 택했다. | ||
|
|
||
| ### 해결 방법 | ||
|
|
||
| `docker/opensql/Dockerfile`을 새로 만들어 `tmaxopensql/postgres:14.6` 위에 pgvector 0.8.0을 소스에서 빌드해 설치했다. | ||
| 기존 이미지 안에 `gcc`, `make`, `pg_config`가 모두 있어서 별도 도구 설치 없이 컴파일이 가능했다. | ||
|
|
||
| ```dockerfile | ||
| FROM tmaxopensql/postgres:14.6 | ||
|
|
||
| USER root | ||
|
|
||
| RUN curl -L https://github.com/pgvector/pgvector/archive/refs/tags/v0.8.0.tar.gz \ | ||
| -o /tmp/pgvector.tar.gz \ | ||
| && tar -xzf /tmp/pgvector.tar.gz -C /tmp \ | ||
| && cd /tmp/pgvector-0.8.0 \ | ||
| && make PG_CONFIG=/usr/pgsql-14/bin/pg_config \ | ||
| && make install PG_CONFIG=/usr/pgsql-14/bin/pg_config \ | ||
| && rm -rf /tmp/pgvector.tar.gz /tmp/pgvector-0.8.0 | ||
|
|
||
| COPY init-and-start.sh /usr/local/bin/init-and-start.sh | ||
| RUN chmod +x /usr/local/bin/init-and-start.sh | ||
|
|
||
| CMD ["/usr/local/bin/init-and-start.sh"] | ||
| ``` | ||
|
|
||
| ### init-and-start.sh | ||
|
|
||
| OpenSQL entrypoint는 Ansible을 실행해 PostgreSQL을 초기화한다. | ||
| Ansible 완료 후 CMD로 지정한 `init-and-start.sh`가 실행되며 아래 작업을 수행한다. | ||
|
|
||
| 1. postgres를 로컬 소켓으로 임시 시작 | ||
| 2. `docgrid` DB 생성 (없으면) | ||
| 3. `docgrid` DB에 `vector` 확장 활성화 | ||
| 4. 외부 TCP 접속 허용을 위해 `pg_hba.conf`에 항목 추가 | ||
| 5. postgres 정지 후 foreground로 재시작 | ||
|
|
||
| ```bash | ||
| pg_ctl start -D "$PGDATA" -l /tmp/pg_init.log -o "-h ''" -w | ||
|
|
||
| psql -d postgres -tc "SELECT 1 FROM pg_database WHERE datname = 'docgrid'" | grep -q 1 \ | ||
| || psql -d postgres -c "CREATE DATABASE docgrid OWNER docgrid;" | ||
|
|
||
| psql -d docgrid -c "CREATE EXTENSION IF NOT EXISTS vector;" | ||
|
|
||
| grep -qxF "host all all 0.0.0.0/0 trust" "$PGDATA/pg_hba.conf" \ | ||
| || echo "host all all 0.0.0.0/0 trust" >> "$PGDATA/pg_hba.conf" | ||
|
|
||
| pg_ctl stop -D "$PGDATA" -m fast -w | ||
| exec postgres | ||
| ``` | ||
|
|
||
| ### 트레이드오프 — pg_hba.conf 규칙 범위 | ||
|
|
||
| 현재 `host all all 0.0.0.0/0 trust`로 설정되어 있다. | ||
| 로컬 개발 환경에서는 문제없지만, 포트가 외부 네트워크에 노출되는 상황(공용 와이파이, 시연 환경 등)에서는 인증 없이 접속이 가능해질 수 있다. | ||
| 시연 전에는 `192.168.65.1/32` 또는 실제 Docker 브리지 대역으로 범위를 좁히는 것을 권장한다. | ||
|
|
||
| ### 참고 — Ansible 재실행 동작 | ||
|
|
||
| OpenSQL entrypoint는 컨테이너가 새로 생성될 때마다 Ansible을 실행한다 (볼륨 유지 여부와 무관). | ||
| 이미 초기화된 경우 대부분의 태스크가 skip되지만 전체 플레이가 돌기 때문에 약 5분이 소요된다. | ||
| `docker compose build` 후 재시작 시에도 동일하게 발생하므로, Dockerfile 변경은 확실한 사항만 모아서 한 번에 반영하는 것이 좋다. | ||
|
|
||
| --- | ||
|
|
||
| ## 2. 임베딩 서버 구축 | ||
|
|
||
| ### 설계 배경 | ||
|
|
||
| pgvector 기반 벡터 검색을 위해 텍스트를 1024차원 벡터로 변환하는 임베딩 서버가 필요했다. | ||
| Spring Boot에서 직접 모델을 실행하기 어렵기 때문에 Python 서버를 별도 서비스로 분리하고 HTTP로 통신하는 구조를 채택했다. | ||
|
|
||
| 모델은 **BAAI/bge-m3**를 사용한다. 다국어(한국어 포함) 지원, 1024차원 dense vector 출력, pgvector HNSW 인덱스와의 궁합이 선택 이유다. | ||
|
|
||
| ### 구현 구조 | ||
|
|
||
| ``` | ||
| embedding-server/ | ||
| ├── main.py # FastAPI 앱 | ||
| ├── requirements.txt # 의존성 (버전 고정) | ||
| └── Dockerfile # python:3.11-slim 기반 | ||
| ``` | ||
|
|
||
| Spring Boot와 같은 `docgrid-local` Docker 네트워크에 올라가며, `http://docgrid-embedding:8000`으로 통신한다. | ||
|
|
||
| ### API 명세 | ||
|
|
||
| #### GET /health | ||
|
|
||
| 서버 및 모델 로드 상태 확인. | ||
|
|
||
| **Response 200** | ||
| ```json | ||
| { "status": "ok" } | ||
| ``` | ||
|
|
||
| **Response 503** — 모델 미로드 시 | ||
| ```json | ||
| { "detail": "Model not loaded" } | ||
| ``` | ||
|
|
||
| #### POST /embed | ||
|
|
||
| 텍스트를 1024차원 벡터로 변환. | ||
|
|
||
| **Request** | ||
| ```json | ||
| { "text": "검색할 텍스트" } | ||
| ``` | ||
|
|
||
| **Response 200** | ||
| ```json | ||
| { "vector": [0.012, -0.034, ..., 0.087] } | ||
| ``` | ||
|
|
||
| **Response 503** — 모델 미로드 시 | ||
| ```json | ||
| { "detail": "Model not loaded" } | ||
| ``` | ||
|
|
||
| ### 주요 설계 결정 | ||
|
|
||
| #### FastAPI 선택 | ||
|
|
||
| async 기반으로 Spring Boot에서 동시 요청이 들어올 때 안정적으로 처리한다. | ||
| Swagger UI가 자동 생성되어 별도 설정 없이 `http://localhost:8000/docs`에서 확인 가능하다. | ||
|
|
||
| #### 모델 런타임 다운로드 + 볼륨 캐시 | ||
|
|
||
| 빌드 시 모델을 이미지에 포함하면 이미지 크기가 3GB 이상 커진다. | ||
| 대신 첫 컨테이너 시작 시 HuggingFace에서 다운로드하고 `huggingface-cache` Docker 볼륨에 캐시한다. | ||
| 두 번째 실행부터는 볼륨에서 로드하므로 다운로드 없이 빠르게 뜬다. | ||
|
|
||
| ```yaml | ||
| volumes: | ||
| - huggingface-cache:/root/.cache/huggingface | ||
| ``` | ||
|
|
||
| #### 의존성 버전 고정 | ||
|
|
||
| `FlagEmbedding==1.2.11`이 최신 `transformers`(4.47+)를 끌어오면 `torch 2.4.1`과 DTensor import 충돌이 발생한다. | ||
| `transformers==4.44.2`로 고정해서 해결했다. | ||
|
|
||
| `peft` 패키지는 `FlagEmbedding`의 reranker 모듈이 의존하지만 자동 설치되지 않아 명시적으로 추가했다. | ||
|
|
||
| ``` | ||
| FlagEmbedding==1.2.11 | ||
| torch==2.4.1 | ||
| transformers==4.44.2 | ||
| peft==0.12.0 | ||
| ``` | ||
|
|
||
| #### healthcheck — curl 대신 python3 urllib | ||
|
|
||
| `python:3.11-slim`에는 `curl`이 없다. | ||
| 별도 패키지 설치 없이 Python 표준 라이브러리로 healthcheck를 구현했다. | ||
|
|
||
| ```yaml | ||
| test: ["CMD", "python3", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/health')"] | ||
| ``` | ||
|
|
||
| ### 로컬 실행 방법 | ||
|
|
||
| ```bash | ||
| docker compose build embedding-server | ||
| docker compose up -d embedding-server | ||
| ``` | ||
|
|
||
| - 첫 실행 시 bge-m3 모델 다운로드로 약 10~15분 소요 (약 3GB) | ||
| - `docker logs -f docgrid-embedding` 으로 진행 상태 확인 | ||
| - `Uvicorn running on http://0.0.0.0:8000` 로그가 뜨면 준비 완료 | ||
| - 이후 재시작은 볼륨 캐시에서 로드하므로 빠름 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,10 @@ | ||
| FROM python:3.11-slim | ||
|
|
||
| WORKDIR /app | ||
|
|
||
| COPY requirements.txt . | ||
| RUN pip install --no-cache-dir -r requirements.txt | ||
|
|
||
| COPY main.py . | ||
|
|
||
| CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"] |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,53 @@ | ||
| import threading | ||
| from contextlib import asynccontextmanager | ||
| from fastapi import FastAPI, HTTPException | ||
| from pydantic import BaseModel | ||
| from FlagEmbedding import BGEM3FlagModel | ||
| import logging | ||
|
|
||
| logger = logging.getLogger(__name__) | ||
|
|
||
| model: BGEM3FlagModel | None = None | ||
|
|
||
|
|
||
| def _load_model(): | ||
| global model | ||
| logger.info("Loading BAAI/bge-m3 model...") | ||
| model = BGEM3FlagModel("BAAI/bge-m3", use_fp16=True) | ||
| logger.info("Model loaded.") | ||
|
|
||
|
|
||
| @asynccontextmanager | ||
| async def lifespan(app: FastAPI): | ||
| thread = threading.Thread(target=_load_model, daemon=True) | ||
| thread.start() | ||
| yield | ||
| global model | ||
| model = None | ||
|
|
||
|
|
||
| app = FastAPI(lifespan=lifespan) | ||
|
|
||
|
|
||
| class EmbedRequest(BaseModel): | ||
| text: str | ||
|
|
||
|
|
||
| class EmbedResponse(BaseModel): | ||
| vector: list[float] | ||
|
|
||
|
|
||
| @app.get("/health") | ||
| def health(): | ||
| if model is None: | ||
| raise HTTPException(status_code=503, detail="Model not loaded") | ||
| return {"status": "ok"} | ||
|
|
||
|
|
||
| @app.post("/embed", response_model=EmbedResponse) | ||
| def embed(req: EmbedRequest): | ||
| if model is None: | ||
| raise HTTPException(status_code=503, detail="Model not loaded") | ||
| result = model.encode([req.text], batch_size=1, max_length=8192) | ||
| vector = result["dense_vecs"][0].tolist() | ||
| return EmbedResponse(vector=vector) | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,7 @@ | ||
| fastapi==0.115.0 | ||
| uvicorn==0.30.6 | ||
| FlagEmbedding==1.2.11 | ||
| torch==2.4.1 | ||
| transformers==4.44.2 | ||
| peft==0.12.0 | ||
| numpy==1.26.4 |
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
PR 목표에 명시된
normalize_embeddings=True적용이 누락되었습니다.PR 요약 및 목표(PR objectives)에는
normalize_embeddings=True를 적용한다고 기재되어 있으나,model.encode()호출에는 해당 옵션이 반영되어 있지 않습니다. 의도적으로 제외하신 것인지, 아니면 추가해야 하는지 확인 부탁드립니다.💡 옵션 추가 제안
만약 라이브러리에서 해당 인자를 지원한다면 다음과 같이 추가할 수 있습니다:
📝 Committable suggestion
🤖 Prompt for AI Agents