Companies & labs
Straight from the people who built it
Every post here is an organisation publishing about its own work: a lab's newsroom, a research blog, a paper. No intermediary, no summary of a summary, and nothing this deck had to interpret. It is the shortest distance between you and the source.
Newest first
- PrimaryNVIDIA21 Sep, 1:45 PM CDT
Benchmarking LLM Inference at Scale with AIPerf
You’re deploying a model on a system. It starts up, prompts are getting responses. Now the hard question: Is this fast? Your instincts might lead you to send...
- PrimaryAWS Machine Learning21 Sep, 1:30 PM CDT
xAI’s Grok 4.6 is now available in Amazon Bedrock
xAI's Grok 4.6 is now available in Amazon Bedrock: a frontier model for long-running agents, coding, and knowledge work, with a 500K token context window and four reasoning effort levels. It runs on both the bedrock-mantle and bedrock-runtime endpoints, with C
- PrimaryAWS Machine Learning21 Sep, 11:34 AM CDT
Run Positron on Amazon SageMaker AI for data science workflows
Positron, Posit's IDE for data science, now runs on Amazon SageMaker AI. This post shows how a data scientist explores an Amazon Athena table, validates features in R, trains an XGBoost model in Python, deploys a real-time SageMaker AI endpoint, and reports re
- PrimaryAWS Machine Learning21 Sep, 11:27 AM CDT
How Benchling secured multi-tenant AI agents with Amazon Bedrock AgentCore
Learn how Benchling built a defense-in-depth security architecture to run untrusted, AI agent-generated scientific code across thousands of life sciences tenants using Amazon Bedrock AgentCore Code Interpreter in VPC mode, combined with Amazon Route 53 Resolve
- PrimaryAWS Machine Learning21 Sep, 11:24 AM CDT
Reducing medical claims review time with AI on AWS: The EXL Medical IDP solution
EXL built an AI-powered Medical intelligent document processing (IDP) solution on AWS, combining IDP with domain-specific large language models on Amazon SageMaker and Amazon Bedrock to extract, summarize, and query medical records at enterprise scale and cut
- PrimaryHugging Face21 Sep, 8:44 AM CDT
Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem
- PrimaryOpenAI21 Sep, 7:00 AM CDT
Higgsfield AI ships new video features in a day with GPT-6 Astra
With GPT-6 Astra, Higgsfield AI makes video ad creation easier for small businesses and brings new creative tools to market faster.
- PrimaryOpenAI21 Sep, 7:00 AM CDT
Advisory Group on Mathematics and Artificial Intelligence
OpenAI is working with an independent Advisory Group on Mathematics and Artificial Intelligence to guide the review and communication of emerging AI results.
- PrimaryOpenAI21 Sep, 5:00 AM CDT
Building standards for the next phase of AI
OpenAI outlines a path to shared global AI standards, calling for coordinated evaluation, reporting, and governance to improve safety.
- PrimaryOpenAI21 Sep, 2:00 AM CDT
Expanding OpenAI Academy with new learning paths
Explore new OpenAI Academy learning paths for employees, developers, leaders, educators, and students to build and demonstrate practical AI skills.
- PrimaryOpenAI20 Sep, 7:00 PM CDT
How V7 gives AI agents institutional memory
Using GPT-5.6, V7 turns scattered company files into context agents can use to complete complex, source-linked work.
- PrimaryAWS Machine Learning18 Sep, 3:52 PM CDT
Amazon SageMaker Inference: 2026 year-to-date launches in review
Amazon SageMaker AI shipped 13 inference launches in year-to-date across two deployment paths: fully managed endpoints and Amazon SageMaker HyperPod Inference. This post reviews each launch, from inference recommendations and capacity-aware instance pools to t
- PrimaryarXiv18 Sep, 12:59 PM CDT
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen fron
- PrimaryarXiv18 Sep, 12:55 PM CDT
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, l
- PrimaryarXiv18 Sep, 12:55 PM CDT
Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw
Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize. Using Value Sensitive Design, we analyzed, with LLM assistance, 73,093 first-person Reddit posts about using Op
- PrimaryarXiv18 Sep, 12:54 PM CDT
Advances in large-scale neural recording have made it possible to collect data across many animals and distributed brain regions, raising the question of whether this scale can be exploited to learn general-purpose neural representations transferable across di
- PrimaryarXiv18 Sep, 12:48 PM CDT
Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention
Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility): confident-failure reduction is achievable if and onl
- PrimaryarXiv18 Sep, 12:46 PM CDT
Benchmarking World Models for Continual Learning on Compositional Tasks
A desirable property of a world model is the ability to learn continually across tasks, adapting to new environments without forgetting what the agent has already learnt. In particular, the ability to retain and reuse knowledge obtained from prior experiences
- PrimaryarXiv18 Sep, 12:34 PM CDT
Memory systems for large language models have focused predominantly on efficient retrieval, whereas the decision of whether retrieved memories should be trusted has received comparatively little attention. When the memory store contains conflicting positions,
- PrimaryarXiv18 Sep, 12:33 PM CDT
$λ$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource
Reinforcement learning is increasingly used to align image generators with reward signals, and Flow-GRPO recently extended this paradigm to flow-matching models by treating the denoising sampler as a stochastic policy that can be optimized from reward feedback
- PrimaryarXiv18 Sep, 12:31 PM CDT
Gricea: An Open Science Platform for Conversational AI Research
We need studies on conversational AI (CAI) at scale to understand human behavior and shape CAI design. However, fragmented reporting of systems and study configurations hinders replication, extension, and knowledge accumulation. We present Gricea, an open-scie
- PrimaryarXiv18 Sep, 12:30 PM CDT
We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Quranic benchmarks center on general question answering and semantic retrieval, without probing specific linguist
- PrimaryarXiv18 Sep, 12:08 PM CDT
COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules
Every multiparameter persistence vectorization we know of carries a one-sided Lipschitz upper bound and nothing below it: without a lower gauge there is no sense in which the features are faithful, and no per-prediction guarantee can be built on them. This pap
- PrimaryarXiv18 Sep, 12:06 PM CDT
DiaVLo: Diagnosing Behaviours of Vision-Language Models
Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their reliable deployment. Yet, methods that iden
- PrimaryarXiv18 Sep, 12:03 PM CDT
Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention
Gating the value pathway of attention reportedly improves language model pretraining, and prior studies disagree on why. We argue and provide experimental evidence that such gates supply two different things that softmax attention lacks: abstention and noise f
- PrimaryarXiv18 Sep, 12:00 PM CDT
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Real digital work requires both, interleaved rather than stacked end to end. We study hybrid CUAs that autonomo
- PrimaryarXiv18 Sep, 11:57 AM CDT
Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents
LLM agents in social simulation revise their opinions implicitly, in context: how open an agent is to persuasion can neither be specified nor verified, and collective outcomes inherit the model's training prior. We introduce Bayesian Chronicle Agents (BCA), a
- PrimaryarXiv18 Sep, 11:57 AM CDT
A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
Large language models can hold knowledge they do not report. A model may sandbag on a capability evaluation, or answer against what it internally knows, and its outputs alone cannot tell whether it is hiding an answer or simply does not have one. We borrow the
- PrimaryarXiv18 Sep, 11:56 AM CDT
The prediction of critical heat flux (CHF), a key safety-related quantity in nuclear thermal hydraulics, remains an important challenge due to its direct relationship with fuel performance and reactor safety. Recent studies have demonstrated that relative to t
- PrimaryAWS Machine Learning18 Sep, 11:52 AM CDT
Introducing Kimi K3 on Amazon Bedrock
Kimi K3 from Moonshot AI is now available on Amazon Bedrock, giving you a powerful new open-weight option for coding and knowledge work. It offers native vision, a 1-million-token context window, and explicit prompt caching to reduce latency and input costs.
- PrimaryarXiv18 Sep, 11:27 AM CDT
NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent
- PrimaryarXiv cs.LG18 Sep, 11:18 AM CDT
Schedule optimization for tau-leaping in masked discrete diffusion
Masked discrete diffusion models are commonly accelerated using the so-called tau-leaping discretization method, which reveals several coordinates in parallel at each sampling step. The sampler replaces the joint conditional law of each revealed block by a pro
- PrimaryarXiv cs.LG18 Sep, 11:11 AM CDT
RACER: Role-Aligned Competence Estimation for Human-AI Routing
Learning to defer asks a predictive system when to act autonomously and when to defer to a human expert. Population-adaptive deferral extends this problem to unseen experts using a small context set of expert behavior. Neural context encoders such as L2D-Pop c
- PrimaryarXiv cs.LG18 Sep, 11:03 AM CDT
Urban transportation networks present complex optimization challenges spanning calibration of high-fidelity simulators and real-time operational control. This paper presents a shared latent-space framework that connects simulator calibration and reinforcement
- PrimaryarXiv cs.LG18 Sep, 11:00 AM CDT
Guiding Agents of Quantum Games to Equilibrium using Matrix Exponential Fixed-Point Iteration
In recent years, quantum game theory has gained significant attention as a framework for studying decision-making in multi-agent systems using quantum principles. However, computing equilibrium strategies is challenging because the dimension of the joint Hilbe
- PrimaryarXiv cs.LG18 Sep, 10:56 AM CDT
End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery
The importance of deep neural networks (DNNs) is widely recognized, and the parameters obtained through training are regarded as valuable assets. Recently, attacks that extract these parameters using only oracle queries to a DNN have been actively studied at I
- PrimaryarXiv cs.LG18 Sep, 10:48 AM CDT
Kinks vs. Smoothness: Identifiability of Real Analytic nICA for Laplace-like Sources
Many machine learning systems try to explain complex data - like images or financial time series - in terms of hidden, independent factors that generated them. Recovering the true underlying factors, rather than some scrambled version of them, is the central c
- PrimaryAWS Machine Learning18 Sep, 10:38 AM CDT
Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime
Migrate a multi-model healthcare AI agent from self-managed Amazon ECS with AWS Fargate to Amazon Bedrock AgentCore runtime, preserving triple-model orchestration and vector-enhanced knowledge retrieval while reducing infrastructure management. The framework-a
- PrimaryarXiv cs.LG18 Sep, 10:34 AM CDT
Riemannian Simultaneous Inference for Tangent Vector Field Regression
We consider nonparametric tangent vector field regression on a Riemannian manifold without boundary. Because responses at different points lie in different tangent spaces, the proposed kernel estimator first parallel transports nearby responses to the target t
- PrimaryarXiv cs.LG18 Sep, 10:34 AM CDT
Beyond Kinematics: Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning
In this work, we conduct a systematic comparison of two state-of-the-art motion-imitation reinforcement learning (MIRL) pipelines, one built on SCONE/HyFyDy and one built on MuJoCo/MyoSim. HyFyDy emphasizes physiological realism through detailed musculotendon
- PrimaryarXiv cs.LG18 Sep, 10:33 AM CDT
Counterfactual simulation with a clinical world model means fixing a patient's history, changing the treatment, and reading off the predicted response. Doing so requires deciding what counts as one intervention. In clinical settings, interventions are document
- PrimaryAWS Machine Learning18 Sep, 10:31 AM CDT
The new AgentCore runtime: Elastic, optimized, and consistently fast starts
Today we are announcing the new AgentCore runtime, a capability of Amazon Bedrock AgentCore built for the speed, flexibility, and cost efficiency that production agents demand. It reclaims memory as sessions release it and delivers consistent cold starts regar
- PrimaryarXiv cs.LG18 Sep, 10:26 AM CDT
ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment
Best-of-$n$ (BoN) sampling is a simple yet effective inference-time alignment method, but hard maximization provides only coarse control over the trade-off between reward and distribution shift. Soft Best-of-$n$ (Verdun et al. 2025) provides smoother control a
- PrimaryAWS Machine Learning18 Sep, 10:25 AM CDT
Deploy Hugging Face models on Amazon SageMaker AI with coding agents
Deploy production-ready Hugging Face models on Amazon SageMaker AI using six open-source agent skills. Point a coding agent at a model and get back a real-time endpoint with the right serving container, autoscaling, Amazon CloudWatch alarms, and a verified tea
- PrimaryarXiv cs.LG18 Sep, 10:22 AM CDT
LLMs as Feature Engineers for Text-and-Tabular Prediction
We introduce an iterative framework that automates the extraction of interpretable, schema-bound categorical features from unstructured text for tabular prediction models. To navigate the feature space, a generator LLM proposes semantic definitions, a separate
- PrimaryarXiv cs.CL18 Sep, 10:16 AM CDT
Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective
Detecting pretraining data in large language models is challenging because high likelihood can reflect either training exposure or strong generalization. In the joint space of prediction loss and predictive entropy, a likelihood-only detector uses a horizontal
- PrimaryarXiv cs.CL18 Sep, 9:53 AM CDT
TrialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization
Nearly 90% of drugs entering clinical development ultimately fail, despite billions of dollars in investment. Pharmaceutical companies therefore rely on clinical development planning (CDP) and probability of technical and regulatory success assessment to antic
- PrimaryarXiv cs.CL18 Sep, 9:52 AM CDT
Do Personality-Tuned LLMs Make Better Social Agents?
LLMs are increasingly used in social simulations for socially interactive agents and robots, offering more flexibility than rule-based systems. However, even though they mimic human behaviour very well, there is a persistent alienness to them. This work invest
- PrimaryarXiv cs.CL18 Sep, 9:41 AM CDT
Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts
Long transcripts are costly inputs for downstream NLP systems and often contain irrelevant context. We study query-conditioned topic localization: predicting the sentence span in a transcript that best addresses a topic-title query. To improve span localizatio
- PrimaryarXiv cs.CL18 Sep, 9:28 AM CDT
RheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree Speculative Decoding
Speculative decoding accelerates LLM inference by drafting multiple tokens in parallel, with tree-based methods further improving efficiency through hierarchical structures. Dynamic-tree methods such as EAGLE-3 perform well under greedy decoding via determinis
- PrimaryarXiv cs.CL18 Sep, 9:06 AM CDT
CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation
Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine them. Prior empirical studies,
- PrimaryarXiv cs.CL18 Sep, 9:04 AM CDT
Most multilingual dysarthria-severity systems either train on a single aetiology-language pair or pool heterogeneous aetiologies into one label space. We test that pooling assumption with four matched HuBERT-base contrastive embedding models under a shared bac
- PrimaryGoogle AI18 Sep, 9:00 AM CDT
New experts join Google’s AI & Economy team
Text "AI & Economy Research Program" all over a green grid background, with the Google G logo in the bottom right corner
- PrimaryarXiv cs.CL18 Sep, 8:24 AM CDT
World Modeling in Transformers
Behavioral failures can make a transformer appear to lack a world model even when it has learned faithful representations of its environment. We demonstrate this in TaxiGPT, a transformer trained on random walks through Manhattan whose failures have been inter
- PrimaryAWS Machine Learning18 Sep, 8:08 AM CDT
Introducing Amazon SageMaker HyperPod Inference Gateway
Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing add-on for Amazon EKS. It uses real-time GPU signals to send each inference request to the best-suited pod, cutting first-token latency by up to 82% with no changes to your m
- PrimaryarXiv cs.CL18 Sep, 7:57 AM CDT
CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords
Chinese social media has generated a vast and continually evolving lexicon of internet buzzwords whose meanings are often non-literal and deeply rooted in local cultural and pragmatic contexts. Existing research has primarily focused on interpreting these buzz
- PrimaryarXiv cs.CL18 Sep, 7:13 AM CDT
The Spoken Wikipedia Presentation Corpus
We present the Spoken Wikipedia Presentation Corpus, an extension of the Spoken Wikipedia Corpora featuring LLM-generated slide decks for multimodal ASR. Slides are created from LLM-segmented sections using a hybrid pipeline that combines LLM-based content pla
- PrimaryarXiv cs.CL18 Sep, 7:09 AM CDT
PRISM-BN: A Controlled Corpus and Benchmark for Text-to-Parameterized Bayesian Network Extraction
Probabilistic Graphical Models (PGMs), especially Bayesian Networks (BNs), expose directed structure and probabilistic parameters, making them natural symbolic targets for neurosymbolic AI. Yet training text-to-parameterized-BN systems requires paired text-to-
- PrimaryarXiv cs.CL18 Sep, 7:08 AM CDT
Accelerating Dense LLMs via L0-regularized Mixture-of-Experts
Large language models (LLMs) achieve strong performance but suffer from slow and costly inference. Existing acceleration methods often lead to noticeable performance degradation, while Mixture-of-Experts (MoE) models require extensive computational resources.
- PrimaryOpenAI18 Sep, 7:00 AM CDT
Introducing the Australian Youth Safety Blueprint
OpenAI introduces the Australian Youth Safety Blueprint, a six-pillar roadmap for safer AI experiences that protect and empower young people.
- PrimaryarXiv cs.CL18 Sep, 6:58 AM CDT
When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap
Activation steering has become a widely used approach for controlling language models during explicit chain-of-thought (CoT) reasoning, motivating its extension to latent CoT. However, we find that steering continuous thoughts produces substantially weaker eff
- PrimaryarXiv cs.CL18 Sep, 6:53 AM CDT
Analysing the Linearity of Linguistic Relations in Language Model Embedding Spaces
We propose a framework to analyse how strongly different linguistic relations are linearly encoded in language model embedding spaces. We formalise linear encoding via a constrained linear approximation over related and unrelated word pairs and apply this to a
- PrimaryTogether AI17 Sep, 7:00 PM CDT
How a global fintech scaled coding agent traffic with Dedicated Model Inference
Inside a global bank's shift to self-serve dedicated inference: how Together's DMI gave engineering teams direct control over scaling, models, and testing.
- PrimaryNVIDIA17 Sep, 3:20 PM CDT
CUDA Toolkit 13.4 Adds Windows on Arm Support and Greater Control over Shared GPUs
Every NVIDIA CUDA Toolkit release adds functionality and performance improvements that help developers get more from NVIDIA GPUs and the broader NVIDIA software...
- PrimaryNVIDIA17 Sep, 2:23 PM CDT
Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine
Mixture of experts (MoE) has become one of the defining architectural trends in large-scale AI model training. DeepSeek, Qwen, and Mixtral are examples of MoE...
- PrimaryAWS Machine Learning17 Sep, 12:55 PM CDT
Reduce time-to-hire for quality candidates with AI-powered Amazon Connect Talent
Amazon Connect Talent is an AI hiring solution built for talent acquisition leaders managing scaled hiring. It delivers AI-led interviews, data-driven assessments, and consistent evaluation, helping recruiters identify strong candidates more efficiently while
- PrimaryNVIDIA17 Sep, 11:21 AM CDT
How to Use AI Agents to Prepare 3D Scenes for Simulation
Agentic AI workflows can be used to prepare and validate digital twins for physical AI systems. Agents can inspect 3D scenes, author simulation-relevant data in...
- PrimaryAWS Machine Learning17 Sep, 10:53 AM CDT
Selecting a vector store for Amazon Bedrock Knowledge Bases
Choosing the right vector store for your Amazon Bedrock Knowledge Bases RAG application affects performance and cost. This post compares Amazon OpenSearch Service, Amazon Aurora PostgreSQL with pgvector, and Amazon S3 Vectors across three RAG use cases, with b
- PrimaryAWS Machine Learning17 Sep, 10:41 AM CDT
A shared agentic platform for Wood Mackenzie, on Amazon Bedrock AgentCore
Wood Mackenzie built APEX, a shared agentic AI platform on Amazon Bedrock AgentCore so every team can ship production agents without rebuilding runtime, identity, observability, and guardrails from scratch. Learn why they chose AgentCore, how APEX Studio opera
- PrimaryAWS Machine Learning17 Sep, 10:36 AM CDT
How MRH Trowe enabled secure self-service AI agents in financial services
Learn how MRH Trowe, one of Germany's leading commercial and industrial insurance brokers, gave about 400 employees secure, self-service access to AI agents in its first month of production - using Strands Agents, Amazon Bedrock AgentCore, and LibreChat to mee
- PrimaryAWS Machine Learning17 Sep, 10:28 AM CDT
Enhancing industrial safety AI with synthetic data on Amazon SageMaker AI
Learn how to build a synthetic data augmentation pipeline on Amazon SageMaker AI and Amazon Rekognition that generates photo-realistic, auto-labeled training images for industrial safety AI. This approach improved person detection by up to 160% without manual
- PrimaryAllen Institute for AI17 Sep, 3:00 AM CDT
What a crowdsourced game revealed about steering Olmo 3
A crowdsourced game built on Olmo 3 showed how people can exploit unexpected model behaviors to stress-test prosocial AI evaluations—and how open access to a model’s internals can help researchers understand why those tests break.
- PrimaryApple Machine Learning16 Sep, 7:00 PM CDT
REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff
A central goal of autonomous reinforcement learning is continuous policy training without external resets. However, existing paradigms largely depend on underlying environmental reversibility, a property absent in real world manipulation, where events such as
- PrimaryOpenAI16 Sep, 7:00 PM CDT
OpenAI for Law brings frontier intelligence for law, custom firm workflows, connected legal data sources, and legal-grade controls for confidential client work.
- PrimaryNVIDIA16 Sep, 4:46 PM CDT
Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding
Alibaba released the model weights for Qwen3.8-Flash-Next as a preview of the upcoming Qwen4 architecture for developers to experiment with and evaluate. It’s...
- PrimaryNVIDIA16 Sep, 4:39 PM CDT
Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72
Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, bringing near-frontier capabilities to the open...
- PrimaryNVIDIA16 Sep, 3:37 PM CDT
TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor
AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answers a single prompt, an agent works through...
- PrimaryAWS Machine Learning16 Sep, 2:00 PM CDT
Improving HCLS AI reasoning with open-source agent skills
AI agents on foundation models often misapply healthcare and life sciences decision frameworks, citing the right guideline but applying it incorrectly. This post shares 38 open-source agent skills across 11 HCLS domains that close this gap, with installation s
- PrimaryAWS Machine Learning16 Sep, 1:59 PM CDT
Fault tolerant distributed training on Amazon EKS using NVRx
Integrate NVIDIA Resiliency Extension (NVRx) into PyTorch FSDP training on Amazon EKS to overlap checkpoint I/O with training and recover from GPU faults in seconds. This post covers async checkpointing, in-process restart, and ft_launcher in-job restart, with
- PrimaryOpenAI16 Sep, 12:00 PM CDT
Our framework for reporting model misalignment
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
How a post gets here
Evidence tier 1, and nothing else. This deck defines that tier as the organisation said it itself, or this is the paper, and the tier is set per source from what the source IS rather than from how good it is. So a lab's own newsroom qualifies and an outlet reporting on that lab does not, however good the reporting. 21 of the deck's sources carry that tier.
Every link went through the reader's door first, the same as every article link on this site: 0 post(s) were dropped because a reader could not open them. And this page has its own feed at companies.xml, carrying the same items under the same permission law as the main feed.
Publishing here: AWS Machine Learning · Allen Institute for AI · Apple Machine Learning · Berkeley BAIR · EleutherAI · Google AI · Google DeepMind · Google Research · Hugging Face · IBM Research · Microsoft Research · Mistral AI · NIST · NVIDIA · OpenAI · Qwen (Alibaba) · Stability AI · Together AI · arXiv · arXiv cs.CL · arXiv cs.LG