Chuyển tới nội dung chính

Bắt đầu nhanh

Deploy model endpoint đầu tiên và thực hiện lời gọi API đầu tiên. Bạn cần một tài khoản FPT AI Factory còn credit — vào Billing để nạp thêm trước khi bắt đầu.


Bước 1: Deploy một model

  1. Ở sidebar bên trái, trong Dedicated Inference, nhấn Model Catalog.
  2. Duyệt catalog và nhấn Deploy trên thẻ model. Để thử nhanh, chọn model nhỏ (ví dụ Mistral 7B).

Dedicated Inference — Model Catalog

  1. Trong modal Deploy Dedicated Endpoint:
    • Nhập Hugging Face repo / URL (ví dụ mistralai/Mistral-7B-Instruct-v0.3).
    • Chọn Region.
    • Chọn chế độ Quick và chọn Size (Small / Medium / Large).
    • Kiểm tra ESTIMATED COST (PER HOUR / BIZ HOURS / 24/7).
  2. Nhấn Deploy endpoint.

Modal Deploy Dedicated Endpoint

Xem hướng dẫn chi tiết: Deploy một model


Bước 2: Chờ endpoint chuyển sang Running

  1. Ở sidebar bên trái, trong Dedicated Inference, nhấn My endpoints.
  2. Tìm endpoint vừa tạo — trạng thái đang là Deploying.
  3. Chờ đến khi trạng thái chuyển sang Running.

Tab My endpoints hiển thị trạng thái các endpoint

Xem hướng dẫn chi tiết: Quản lý endpoint


Bước 3: Lấy API key

  1. Nhấn tab API keys trong Dedicated Inference.
  2. Nhấn + Create new key.
  3. Nhập Name cho key (ví dụ quickstart-test).
  4. Nhấn Generate Key, rồi copy key — key chỉ hiển thị một lần duy nhất.

Hộp thoại Create an API Key

Xem hướng dẫn chi tiết: Tạo API key


Bước 4: Gọi API

Dùng gateway URL lấy từ trang chi tiết endpoint:

curl https://<ddi-gateway-domain>/v1/chat/completions \
-H "Authorization: Bearer sk-vgw-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx" \
-H "Content-Type: application/json" \
-d '{
"model": "mistralai/Mistral-7B-Instruct-v0.3",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 256
}'

Gateway domain nằm ở trang chi tiết endpoint. Lấy API key tại My endpoints → API keys.

Xem hướng dẫn chi tiết: Gọi Inference API


Bước 5: Dừng endpoint khi dùng xong

Để tránh phát sinh phí GPU không cần thiết, dừng endpoint khi không dùng:

  1. Ở sidebar bên trái, trong Dedicated Inference, nhấn My endpoints.
  2. Tìm endpoint và nhấn Pause ở cột Actions.
  3. Xác nhận thao tác.

Xem hướng dẫn chi tiết: Quản lý endpoint