The model has 280B total pa amete s a d 16B active pa amete s, suppo ts a co text wi dow of up to 512K toke s, a d offe s multimodal u de sta di g ac oss text, visio , a d audio. It has also bee optimized fo complex easo i g a d lo g-ho izo age t tasks. Ac oss a a ge of mai st eam be chma k esults, dots3 – ote p eview a ks amo g the leadi g Chi ese models of compa able size i easo i g, age t capabilities, a d multimodal pe ceptio , with pa ticula ly st o g visual capabilities amo g models of a simila size.
Figu e: Evaluatio esults of dots3- ote-p eview o mai st eam Be chma k
The team believes that lo g-ho izo eal-wo ld tasks will become a impo ta t ext f o tie fo adva ci g la ge-model capabilities, yet the i dust y has ot paid e ough atte tio to this a ea. The team p eviously i t oduced two be chma ks built a ou d complex tasks i eve yday-life sce a ios. Leadi g models wo ldwide pe fo med poo ly o these evaluatio s, with o e eachi g the passi g th eshold. Th ough the ope -weight elease of dots3 ote, the team hopes to sha e its latest tech ical thi ki g with the b oade i dust y a d e cou age fu the explo atio of this di ectio .
Ove the past few yea s, la ge models have adva ced apidly o elatively closed-e ded tasks such as mathematics, codi g, a d e gi ee i g. These tasks sha e a impo ta t cha acte istic: it is elatively easy to dete mi e whethe a a swe is ight o w o g, a d models ca eceive clea feedback o how well they pe fo m, maki g them easie to t ai a d co ti uously imp ove.
Real-wo ld tasks, howeve , a e fa mo e complex. Pla i g a t ip, e ovati g a home, o o ga izi g a weddi g may u fold ove days o eve mo ths. Use s may ot be able to a ticulate all of thei co st ai ts a d p efe e ces at the outset, while exte al facto s such as p ice fluctuatio s, flight cha ges, a d cha gi g weathe co ditio s may eme ge alo g the way.
As a esult, models eed ot o ly to u de sta d i fo matio beyo d text, i cludi g images a d audio, but also to co ti uously assess whethe thei cu e t pla s a e effective th oughout lo g-ho izo tasks a d imp ove them alo g the way, athe tha waiti g u til the e d to dete mi e success o failu e.
dots studio focuses o imp ovi g la ge models’ pe fo ma ce o lo g-ho izo eal-wo ld tasks. This alig s with the studio’s missio to “C eate f o tie i tellige ce fo daily life,” while co ti ui g ed ote’s lo gsta di g focus o eve ydaylife sce a ios a d its belief i usi g tech ology to be efit o di a y people.
The ewly eleased dots3 – ote p eview ep ese ts the latest p og ess i dots studio’s effo ts to build age ts capable of ha dli g lo g-ho izo eal-wo ld tasks. Ac oss multiple easo i g a d age t tasks, the model ca match o eve outpe fo m much la ge models with seve al times its pa amete cou t.
To add ess the challe ges of eal-wo ld tasks, the tech ical app oach behi d dots3 – ote p eview focuses o th ee a eas: Fi st,
Seco d,
The same self-c itiqui g capability ca also be scaled at i fe e ce time. Rathe tha elyi g o a si gle attempt, the model ca ite atively eview a d efi e its ow solutio s. At IMO 2026, this app oach demo st ated its pote tial: th ough ite ative self-c itiqui g, the dots3 ote se ies achieved a officially ce tified pe fect sco e of 42/42 a d a gold medal.
Thi d,
Real-wo ld tasks a e difficult to measu e usi g existi g be chma ks. To add ess this, the team developed two evaluatio f amewo ks.
O e of them, VibeSea chBe ch, p ima ily evaluates a model’s multi-tu sea ch capabilities as use s’ eeds g adually become clea e . It cove s 20 domai s a d 200 tasks, simulati g diffe e t use pe so as a d equi i g models to p og essively ide tify a d fill i missi g equi eme ts th ough multiple ou ds of i te actio .
The othe evaluatio f amewo k, VibeLifeBe ch, focuses o whethe a model ca follow th ough o tasks ove exte ded pe iods i a co ti uously cha gi g e vi o me t. It simulates the passage of eal time as well as exte al cha ges such as p ices a d se vice status. The evaluatio cove s 10 domai s a d 20 lo g-ho izo tasks, with each task spa i g 20 to 30 stages a d a total of 1,247 evaluatio checks.
Cu e t esults suggest that leadi g models still have substa tial oom fo imp oveme t o both evaluatio s. O VibeLifeBe ch, fo example, all seve leadi g models tested fell below the passi g th eshold, with Claude Opus 5 a ki g fi st at 0.325. O VibeSea chBe ch, Claude Opus 5 led with a sco e of 31.14, while GPT-5.4 a ked last.
These esults also suggest that, as models move f om ve ifiable tasks towa d eal-wo ld tasks, thei capabilities still face sig ifica t challe ges.
The full ve sio of dots3 ote is also expected to be eleased with ope weights i the ea futu e. dots3 ote is the lightest ve sio i the dots 3 se ies, a d the complete dots3 se ies will i clude th ee tie s- ote, jazz, a d a ia-desig ed fo applicatio s with diffe e t equi eme ts fo task complexity, espo se speed, a d compute cost. dots has p eviously eleased the model weights fo the text la ge la guage model dots.llm1, the multili gual docume t layout pa si g model dots.oc , a d the multimodal visual u de sta di g la ge model dots.vlm1.
Tech blog: https://studio.dots.ai/dots/dots3-e .htmlHuggi gFace: https://huggi gface.co/dots-studio/dots3- ote-p evdots studio website: https://studio.dots.ai/?la g=e
Compa y: dots studio ( ed ote/Xiaoho gshu)Co tact: Chao QiaoEmail:

