今天,我们分享一项早期预览——Ultrafast,这是一个新的服务层级,运行GPT‑5.6 Sol的速度比标准处理快达14倍,首先在OpenAI API中推出。由Cerebras提供支持,Ultrafast每秒可生成多达750个输出令牌,将我们最智能的模型带入每一秒都至关重要的产品和流程中。
借助GPT‑5.6,我们正在推动模型能力的边界,并在堆栈的每一层提升效率。这些改进使先进智能更加经济实惠,并对更多人更有用。到目前为止,获得实时速度通常意味着选择更小或更专门的模型。Ultrafast指向了一个新方向的进展:每秒完成更多有用工作。
当速度不再需要牺牲智能时,AI就能进入业务中最注重时间的环节,新型工作也变为可能。我们已经看到一些令人鼓舞的Ultrafast应用场景:
- 事件响应和可靠性: 当关键系统发生故障时,在故障仍在持续时,分析应用日志、最近的代码变更和工程师报告,以识别可能的原因并帮助准备修复方案。
- 金融研究和安全: 在市场条件仍在变化时,分析市场信号、评估交易并识别可疑活动。
- 客户支持和语音: 实时解决复杂客户问题而不打断对话,即使找到答案需要多个步骤或系统。
- 商务: 在购物者仍在决策时,回答产品问题、检查库存、个性化推荐并解决结账问题,在犹豫变成放弃购物车之前。
- 实时研究和实验: 将以前需要通宵运行的研究转变为交互式工作会话,让团队测试想法、检查结果、调整方法并运行另一个实验,而不打断他们的工作流。
在预览期间,我们正在与一组初始客户合作,以了解这种速度在哪些方面产生最大差异,以及这些经验如何随时间指导我们的产品。如果您的业务需要最高速度的前沿智能,您可以注册以在访问扩展时获得通知。
GPT‑5.6 Sol Ultrafast和标准版本从相同的文本提示中并排构建一个可工作的3D仓库模拟器。
早期客户的体验
我们一直在与一组涵盖编码、商务、金融研究、支持和其他交互式应用的初始公司测试Ultrafast模式下的GPT‑5.6 Sol。从业务流程开始,让我们能够在真实生产环境中研究这些条件。他们的早期工作帮助我们了解速度的数量级变化在哪些方面创造最大价值,以及当模型能够跟上使用者的节奏时,产品会如何变化。我们将利用这些发现来指导容量增长时的部署。
1 / 4
OpenAI如何使用Ultrafast
在OpenAI内部,一组开发人员一直在测试Ultrafast模式下的GPT‑5.6 Sol,以了解哪些工作流受益于能够实时回应的前沿智能。
事件响应是我们团队使用Ultrafast的一个例子。当警报触发时,工程师需要在系统和证据仍在变化时构建准确的图景。团队使用它快速读取日志、分析追踪、综合对话、识别下一步检查,并帮助准备或验证修复方案——所有这些都在极短时间内完成,并具备Sol的智能。它减少了观察信号、测试假设和选择下一步行动之间的延迟,而工程师仍然负责判断和部署。
对于研究,我们的团队使用Ultrafast快速搜索知识来源、查询数据,并跨连接工具快速收集、组织和总结信息。研究中的一个常见工作流是团队成员在夜间启动一批实验,并在早上查看结果。借助Ultrafast,我们看到这个循环正在收紧,转而支持工作日内多次迭代。
由Cerebras提供支持
Ultrafast标志着我们与Cerebras合作将超低延迟推理引入OpenAI平台的下一步。现在,借助Ultrafast模式下的GPT‑5.6 Sol,Cerebras正在支持OpenAI最智能的模型,每秒提供多达750个输出令牌,使企业能够构建响应更快的产品、做出更快的决策,并将强大的AI直接带入他们最苛刻的工作流中。
可用性
Ultrafast模式下的GPT‑5.6 Sol今天以有限预览形式提供给一组精选客户。随着容量增长,我们将扩大访问范围。注册获取更新。
Today, we’re sharing an early look at Ultrafast, a new service tier that runs GPT‑5.6 Sol up to 14× faster than Standard processing, launching first in the OpenAI API. Powered by Cerebras, Ultrafast generates up to 750 output tokens per second, bringing our most intelligent model to products and workflows where every second matters.
With GPT‑5.6, we’re pushing the frontier on what our models can do and making them more efficient across every layer of our stack. Those improvements have made advanced intelligence more affordable and more useful to more people. Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.
When speed no longer requires giving up intelligence, AI can move into the most time-sensitive parts of a business and new kinds of work become possible. We’ve already seen some encouraging scenarios for Ultrafast:
- Incident response and reliability: When a critical system fails, analyze application logs, recent code changes, and engineer reports to identify the likely cause and help prepare a fix while the outage is still unfolding.
- Financial research and security: Analyze market signals, assess transactions, and identify suspicious activity while conditions are still changing.
- Customer support and voice: Resolve complex customer issues in real time without interrupting the conversation, even when finding the answer requires multiple steps or systems.
- Commerce: Answer product questions, check inventory, personalize recommendations, and resolve checkout issues while the shopper is still deciding, before hesitation becomes an abandoned cart.
- Live research and experimentation: Turn research that previously took an overnight run into an interactive working session, letting teams test an idea, examine the results, adjust their approach, and run another experiment without breaking their flow.
During the preview period, we’re working with an initial group of customers to understand where this speed makes the biggest difference, and how those learnings can inform our products over time. If your business requires frontier intelligence at the highest speed, you can sign up to get notified when access expands.
GPT‑5.6 Sol Ultrafast and standard build a working 3D warehouse simulator from the same text prompt, side by side.
What early customers are experiencing
We’ve been testing GPT‑5.6 Sol on Ultrafast mode with an initial group of companies across coding, commerce, financial research, support, and other interactive applications. Starting with business workflows lets us study these conditions in real production environments. Their early work is helping us understand where an order-of-magnitude change in speed creates the most value and how products change when the model can keep pace with the person using it. We will use these findings to guide deployment as capacity grows.
1 of 4
How OpenAI is using Ultrafast
Inside OpenAI, a group of developers has been testing GPT‑5.6 Sol on Ultrafast mode to understand which workflows benefit from frontier intelligence that can answer in real-time.
Incident response is one example where our team is using Ultrafast. When an alert fires, engineers need to build an accurate picture while the system and the evidence are still changing. Teams use it to quickly read logs, analyze traces, synthesize conversations, identify the next checks, and help prepare or validate a fix—all in a fraction of the time with the intelligence of Sol. It reduces the delay between observing a signal, testing a hypothesis, and choosing the next action, while engineers remain responsible for judgment and deployment.
For research, our team uses Ultrafast to rapidly search knowledge sources, query data, and quickly gather, organize, and summarize information across connected tools. A common workflow in research is for our team members to launch a batch of experiments over night, and review the results in the morning. With Ultrafast, we see this loop tightening to support multiple iterations during the workday instead.
Powered by Cerebras
Ultrafast marks the next step in our partnership with Cerebras to bring ultra-low-latency inference to OpenAI’s platform. Now, with GPT‑5.6 Sol on Ultrafast mode, Cerebras is supporting OpenAI’s most intelligent model, delivering up to 750 output tokens per second, enabling businesses to build more responsive products, make faster decisions, and bring powerful AI directly into their most demanding workflows.
Availability
GPT‑5.6 Sol on Ultrafast mode is available in a limited preview today to a select group of customers. We’ll expand access as capacity grows. Sign up for updates.
本文内容采集自官方网站,排版和翻译可能与原页面存在差异。
阅读官方全文