TECHNOLOGY · LLM DEVELOPMENT
AI Experiment· Engineering · StackASI and digital civilization expressions are part of our research and brand roadmap, and do not represent current product performance or certification.
Dual-head culture LLM builds on Qwen3 4B-class open weights with XenLook fine-tuning, evaluation, and hosting. Dialogue (XenLook 4b-Ko) uses curated SFT then full fine-tuning; code (XenLook-Coder) uses QLoRA + micro-DPO. Tuning and TRACK run on NVIDIA DGX Spark (GB10); production vLLM runs on dedicated RTX GPUs. We do not claim foundation pretraining or universal SOTA.
Recipes shared under NDA for technical cooperation only.
LOCAL CULTURE LLM · ON-PREM
Local storage · inference · Vault. Separate from default web on.
Installable companion and privacy as product constraints.
Local cultural LLM = Companion/Vault only. Not a “sovereign AI company” claim. Cross-border AI transfer follows Privacy §8④ and §12 tables.
Companion on-prem →DEV · PROD INFRA
DGX Spark is the tuning/eval workstation. Prod serving is a separate GPU server.
NVIDIA DGX Spark (GB10 Grace Blackwell) — SFT · full-FT · DPO experiments · TRACK-DLG/CODER eval · vLLM smoke
NVIDIA RTX GPU · vLLM · LiteLLM xenlook-4b-ko / xenlook-coder · on · Companion routing
XenLook 4b-Ko: Qwen3 4B-class → SFT → **full fine-tuning** → IF-repair ops path. XenLook-Coder: QLoRA c1 → micro-DPO merge (**not full-FT**). Both are open-weight backbone fine-tuning, not from-scratch pretrain.
NVIDIA Inception program Stage 1 member. No official NVIDIA case study or product endorsement until NVIDIA publishes one.
KOREAN DATA · LINEAGE
We disclose pipeline stages only — starting from public sources. Filter rules, source mix, and golden-set details are shared under NDA only.
AI Hub raw 20M+ rows
Curation · dedup · quality gates ~900K SFT
Current dialogue engine prod mix 24K
AI Hub datasets follow each license and AI Hub policy. Persona 144K SFT is a separate track from Nemotron 7M source (CC BY 4.0). This is not “we built 20M from scratch” — it is public sources → XenLook curation pipeline.
AI policy · data sources →PRESS · SCORES
Self-measured benchmark scores in plain language. Not a global #1 or third-party audit claim.
Dialogue · full-FT
91 pts
Korean dialogue & persona self-benchmark composite — XenLook 4b-Ko
Self-protocol · TRACK-DLG · closed 2026-06-30
Code · QLoRA+DPO
84 pts
Python coding self-benchmark composite (HumanEval + MBPP average). HumanEval representative task: 91 pts.
Self-protocol · inference A2 (best-of-4) · measured 2026-07-08
XenLook 4b-Ko self-benchmark composite 91 pts (closed 2026-06-30). Self-protocol measurement pending third-party audit.
MODEL CARD
Detailed recipes under NDA cooperation only.
Promoted to production · July 2026
Dialogue · full-FT
Korean persona dialogue
Code · QLoRA+DPO
Python EvalPlus codegen
ROADMAP · (DESIGN GOAL)
Dates and metrics may change.
2026 · SHIPPED
Full-FT dialogue + QLoRA+DPO code · DGX Spark tuning · RTX prod
2027+ · (DESIGN GOAL)
Llama 3.2 1B · SmolLM2-1.7B · Gemma 2/3n E2B · Granite 3.1 2B (TBD)
2027+ · (DESIGN GOAL)
TRACK gates first
We do not claim benchmark superiority without published audit evidence.