Skip to content
akshat.katiyar

The reading room

Engineering blogs
worth your time.

Recent posts and AI research from the companies I follow, pulled from their feeds and linked straight to the originals. Nothing is reproduced here, only pointed at.

  • 1301 posts tracked
  • 37 sources
  • Refreshed daily

Sources

1301 posts, newest first

PinterestFrom Activity to Intent: Generating User Journeys with LLMsLin Zhu | Sr. Staff Machine Learning Engineer; Manan Kalra | Machine Learning Engineer II; Byung Eun Jeon | Sr. Machine Learning Engineer; Ye Liu | Staff Machine Learning Engineer; Xiangyi Chen | Sr. Machine Learning Eng…OpenAIAsana cuts model costs 76x in browser tests with GPT-6.1 SolUsing GPT-6 Astra in Codex, Asana made its browser agent 76x cheaper and 5x faster in tests to offer customers more capable models.OpenAISophos cuts threat investigation time by 96% with OpenAI DaybreakDiscover how Sophos uses OpenAI’s Daybreak to cut cyber-threat investigation time by 96% and automate 52% of MDR cases while preserving human oversight.MicrosoftIntroducing the Agent Experience (AX) Practitioner PlaybookAI coding agents now pick the SDK, the version, and the pattern for your developers. The AX Practitioner Playbook is our method for finding out whether they get your technology right, and for fixing the sources they rely…Hugging FaceImpactful scheduling for GPU clustersGitHubHack the World: Why hackathons are still the best place to learn to buildThe barriers to building software have collapsed. Today, anyone can build. Hackathons are a great place to start learning. The post Hack the World: Why hackathons are still the best place to learn to build appeared first…CohereShared vs. Dedicated AI Inference: Embed & RerankCloudflareDeno is joining CloudflareThe Deno team is joining Cloudflare to radically simplify self-hosting Workers and Durable Objects, so developers can use the same primitives in more places.CloudflareIntroducing on-demand CPU and memory profiling with flamegraphs for Workers and Durable ObjectsOn-demand CPU and memory profiling for Cloudflare Workers and Durable Objects is now available. Now you can generate interactive flamegraphs directly in production to quickly track down memory leaks and performance bottl…CloudflareIntroducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flashWe are expanding the Clef decision model family with Clef-omni, natively processing audio, video, images, and text in a single pipeline. We’ve also lowered Clef-flash pricing and boosted Clef inference speeds by up to 2.…SpotifyIntroducing: Spotify Technology. Proven at Spotify, now yours.Introducing Spotify Technology. The post Introducing: Spotify Technology. Proven at Spotify, now yours. appeared first on Spotify Engineering .SalesforceApache Spark Resource Optimization: Lessons From Eight Million Jobs a DayBy Siddharth Sharma, Suvhrajit Basak, Kusum Dhalia, and Nabeel Qaiser. Your compute bill keeps climbing, and your Spark workloads keep succeeding. Successful execution tells you the configuration worked. It does not tell…OpenAIDisrupting AI-enabled “false front” operationsOpenAI disrupted two AI-enabled influence operations that used false-front journalists and a think tank to spread geopolitical messaging.OpenAILegalOn halves Codex costs while maintaining development speedLegalOn cut estimated daily Codex costs by 65% while maintaining development speed. It matched Astra, Sol, and Luna to tasks and managed budgets strategically.OpenAIPollo AI turns creative ideas into campaigns with OpenAIWith GPT-5.6, GPT-6 Astra, and GPT‑Image‑2.5, Pollo AI helps creators turn bold ideas into detailed images and cinematic video ads.OpenAIHow Oracle turns days of work into minutes with ChatGPT and CodexAcross recruiting, engineering, and operations, Oracle turns specialist knowledge into fast, repeatable workflows with ChatGPT Work and Codex.NVIDIABuilding Reliable Data Analytics Agents: Lessons from the KDD CupThe NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a simple idea of making an agent's harness smaller,...NVIDIA5 Steps to Create SimReady Assets for Robotics with Frontier AI ModelsPreparing CAD assets for robotics simulation requires more than converting geometry to OpenUSD: developers must configure and validate materials, collision...Hugging FaceThe model that didn't exist, so you made it yourselfGitLabTrack organization-wide security risk in one dashboardIf you run application security across more than one top-level group, getting a single organization-wide view of your risk has meant manually pulling together data. That is operational work rebuilt in spreadsheets and on…GitHubHow one bug bounty researcher chooses the features they investigateAs we kick off Cybersecurity Awareness Month, the GitHub Bug Bounty team spotlights @vaib25vicky, exploring their methodology, techniques, and experiences hacking on GitHub. The post How one bug bounty researcher chooses…CohereCohere Expands Canadian Footprint with New Ottawa OfficeCloudflareBridging technical depth and usability: The story behind Radar’s redesignWe redesigned Cloudflare Radar to make real-time global traffic and outage data accessible to a wider audience — including journalists, less technical researchers, and everyday users. By introducing an interactive map an…CloudflareBridging technical depth and usability: The story behind Radar’s redesignWe redesigned Cloudflare Radar to make real-time global traffic and outage data accessible to a wider audience — including journalists, less technical researchers, and everyday users. By introducing an interactive map an…arXiv · NLPFrom Sparse Representations to Behavioral Insights for Multimodal Depression AssessmentMultimodal depression assessment offers a promising approach to analyzing behavioral patterns associated with depression. However, existing methods often rely on dense and opaque multimodal representations, making it dif…arXiv · NLPEasy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching missesByte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch. BLT starts a patch where a small model's next-byte entropy is high, so global…arXiv · NLPMemento 3: Model-Based Recursive Self-Improvement through Reflective RulebooksLearning to act in unfamiliar environments requires agents to infer how the world works and revise that understanding as new evidence arrives. Yet limited observations can support multiple world models that explain past…arXiv · NLPDetecting Spin in Clinical Trials with Large Language ModelsSpin in clinical trials includes reporting practices that distort the presentation of results. This is particularly critical in medicine, where spin is present in more than 50% of randomized controlled trials that fail t…arXiv · NLPGRPODropout: Less is More for Online Reinforcement Learning RolloutsReinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further i…arXiv · NLPForms of LLM-Integrated Applications from LLM-Chats to Autonomous AI Agent SystemLarge language models (LLMs) are increasingly embedded as components in software systems, marketed under labels such as chatbot, copilot, retrieval-augmented generation, workflow, coding agent and AI agent. Whether these…arXiv · NLPCan Decision Models Understand Stance? Evaluating Jev Against General-Purpose LLMsStance detection requires identifying an author's attitude toward a given target, sometimes based on conversational context. Jev, a specialized decision model designed for structured decision-making, offers an alternativ…arXiv · NLPNot Every Change Is Necessary: Recoverable Drift in Large Language Model UnlearningMachine unlearning in large language models aims to remove unwanted knowledge while preserving the model's remaining capabilities. Although existing methods use retention objectives or restrict where edits occur, achievi…arXiv · NLPEvent-Centric Memory with Query-Aware Graph Augmentation for Long-Term Conversational AgentsFor persistent and personalized conversational agents, memory systems can enable them to remember, update, and reason over long histories by storing past interactions and retrieving relevant information. Existing memory…arXiv · NLPProject Greenhouse: Progress Toward Fully Open and Sovereign Agentic SearchProject Greenhouse represents our exploration of a simple thesis: We believe that it is possible to build fully open and sovereign models for agentic search with only modest computational resources. As a first milestone,…arXiv · NLPWhen History Helps and Hurts: Selective History Use across Multimodal TurnsReliable multimodal interaction depends on selective use of conversational history: an earlier question may remain relevant while its previous answer is outdated, whereas a current request may depend on historical eviden…arXiv · NLPMiMo-V2.6: Scaling Reinforcement Learning Towards Self-ImprovementReinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-V2.6 series, an omni-modal family that pushes the frontier of m…arXiv · NLPMindFlow: Mind Supernet Powered Thinking Flows for Research Idea InnovationResearch idea innovation is a fundamental engine of scientific progress, yet it remains difficult to generate and evaluate in a scalable and controllable way. This challenge lies in its inherently open-ended and multi-ob…arXiv · NLPSpecialized Decision Models vs. General-Purpose LLMs: Benchmarking Jev Across Knowledge, Reasoning, and Multilingual TasksJev is a "System One" model that returns a choice among given options instead of generating text. We study how such a specialized decision model compares with general-purpose large language models (LLMs). We evaluate Jev…arXiv · NLPDataVista: Diagnosing Multimodal LLMs on Data Video UnderstandingData video is a media form that integrates data visualization with video narrative, widely adopted in news reporting and business analysis. Compared with general video understanding, data video understanding places great…arXiv · NLPAgentic-TTT: Training test-time policy for test-time trainingTest-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems. By turning de…arXiv · NLPExamining Social Attribution in LLM Reasoning: A Theory-Guided Probing MethodologyLarge language models (LLMs) are increasingly deployed in sociotechnical systems where social attribution, the reasoning process attributing external events to the causes and reasons of agents' social behaviors, plays a…arXiv · NLPInterviewPlayground: A Simulation Environment for Evaluating AI InterviewersIncreasingly, AI interviewers are being developed to elicit open-ended responses in applications like market research, public polling, preference elicitation, and social science research. However, evaluating AI interview…arXiv · NLPNatural Language to First-Order Logic LLM-based AutoformalizationLarge Language Models (LLMs) have renewed interest in autoformalization. Yet, when First-Order Logic (FOL) is considered as the target formalism, the field still lacks a unified task formulation and a systematic survey.…arXiv · NLPWhen Should Agents Think? Adaptive Reasoning via Cross-Turn EstimationLarge language model (LLM)-based agents have demonstrated strong capabilities on complex tasks. They typically perform reasoning before each action throughout an interaction trajectory. However, reasoning may not be nece…arXiv · NLPILM: An AI-Powered Storytelling Educational ToolDigital technologies have made Islamic narratives more accessible, but existing platforms provide limited support for structured learning and comprehension of these stories, particularly in Arabic and multilingual settin…arXiv · NLPAll Verdicts are Not Equal: Rethinking LLM Judge ReliabilityLLM-as-a-Judge is the standard paradigm for NLP evaluation, yet its systemic reliability remains poorly understood despite being widely treated as a deterministic ground truth. We present a comprehensive reliability audi…arXiv · NLPA persistent accuracy ceiling in automated verbal deception detectionAutomated methods have been proposed to overcome the limitations of human verbal deception detection, but evidence remains fragmented across disciplines. We systematically reviewed 25 years of research (289 reports, 6,13…arXiv · NLPRehearse Everything, Remember Nothing: Attic-KV Rehearses What Will Be ReadMany key-value (KV) caches are compressed before anyone knows what will be asked of them: a document cached for retrieval, a prompt prefix shared across requests, the memory of a long conversation. The prevailing approac…arXiv · NLPLanguage-Specific Effects of Tokenizer Choice in Multilingual Language ModelsTokenizer choice affects multilingual language modeling, but vocabulary capacity is finite and vocabulary size is often constrained: improving representation for some languages often comes at the expense of others. We th…arXiv · NLPWhen KL Regularization Misfires in Group Policy OptimizationWhy does removing reference-policy KL regularization sometimes improve group policy optimization? This motivates studying how reference-policy information should enter group-relative updates. We analyze seven potential f…arXiv · NLPSteerablePlex: Can We Steer Full-Duplex Models?Full-duplex speech models can listen and speak simultaneously, enabling natural interaction, but become increasingly difficult to control as the conversation history grows. When used as user simulators, this lack of cont…arXiv · NLPSciTBERT: A family of chronologically consistent language models for scientific and technological language processingPre-trained transformer models are increasingly being used to study scientific and technological progress. Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression…arXiv · NLPDiffuPlex: Accelerating Full-Duplex Spoken Dialog Models via Rolling Masked DiffusionRecent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame. We introduce DiffuPlex, a rolling mask…arXiv · NLPLanguage Models as AI Research World ModelsAI research agents automate the cycle of proposing, implementing, and evaluating experiments, opening a path toward recursive self-improvement. Yet their ability to propose experiments outpaces their capacity to execute…arXiv · NLPTokenRouter: Efficient Serving System for Token-Level LLM RoutingLarge language model (LLM) routing distributes inference work across different models, advancing the cost-quality Pareto frontier of LLM serving. While coarse-grained routing at the session or query level has been widely…arXiv · NLPNativeScope: Relation-Localized Retrieval over Native Topology with a Correct AnchorDense retrieval usually ranks text chunks by their semantic similarity to a question. This ignores structure that many data systems already store, including section membership, session boundaries, and native order. We pr…arXiv · NLPEgoVoice: Proactive Spoken Assistance from Egocentric Multimodal StreamsWearable augmented reality (AR) assistants are moving toward continuous real-world interaction, where they perceive the user's activity through first-person video and audio and provide timely spoken guidance without bein…arXiv · NLPHarnessSQL: Harness-Native Training for SQL Agents in Realistic Database EnvironmentsText-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executin…arXiv · NLPVerdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance SystemsLarge language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given. We test this directly across five models and 20 regulatory and platform-policy domains: de…arXiv · NLPSparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM InferenceThe memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency. Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution t…arXiv · NLPVFold: Symmetry-Aware Cross-Layer Value Cache CompressionWhile caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths. One solution is to compress this memory by exploiting inter-layer cache s…arXiv · NLPCan AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent CompetitionAdversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging. AI agents offer an alternative by turning game experie…arXiv · NLPOvercoming Prior Barriers: Supervised Fine-Tuning under Long-Tail DistributionSupervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support. Frequent concepts are more like…arXiv · NLPAccurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge ConflictWhen retrieved evidence contradicts an agent's prior beliefs, does it revise its answer, acknowledge uncertainty, or persist with an incorrect conclusion? Existing evaluations of agentic systems focus primarily on task s…arXiv · NLPCited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought FaithfulnessLarge language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we s…arXiv · NLPWhich Skill to Distill? SGUID: Selecting a Compact Skill Bank for Model-Skill Co-EvolutionSkills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026). Prior work retrieves skills from a bank by semantic relevance, then uses them as inference-…arXiv · NLPOnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal TransportAgents are deployed in applications from trip planners and stock trading to IT incident triage. In most cases, LLM agents work autonomously with minimal rule-based safeguarding, leading to cost and safety issues from irr…arXiv · NLPLatent Core Tokenizer: Compress, but MeaningfullyTokenizers are commonly optimized for compression, but a compact vocabulary does not necessarily distribute its capacity evenly across languages. We introduce the Latent Core Tokenizer (LCT), a language-agnostic approach…arXiv · NLPLong Text to Predictive Features: LLM-Guided Blockwise Feature Engineering via Executable Program SearchIndustrial risk-control systems typically rely on structured-data models for efficient prediction, yet substantial valuable information remains embedded in unstructured long text. Extracting this information through manu…arXiv · NLPSpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language ModelsExisting spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene fro…arXiv · NLPViSkill: Reinforcing VLM Agents with Evolving Visual-Native SkillsSkill-augmented agents improve sample efficiency by distilling successful trajectories into reusable strategies. Yet most existing approaches remain text-centric, linearizing spatial layouts and action-state corresponden…arXiv · NLPPredicting Alignment Generalization with Value RepresentationsLLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highl…arXiv · NLPWOVEN: Weaving Visual World Modeling into Multimodal LLMsMultimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether thi…arXiv · NLPFastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?Streaming Video Large Language Models (VLMs) enable continuous video understanding, yet existing benchmarks focus on low-dynamic scenarios. Under bounded context budgets, models must balance temporal history, spatial res…arXiv · MLScalable Hierarchical Graph Generation via Soft Community StructureGenerating large attributed graphs requires reproducing the topology, generating attributes jointly with the structure, and remaining scalable. Many real-world graphs exist as a single large graph, so a generative model…arXiv · MLIs Real-World Training Data Necessary for Generalist Graph Anomaly Detection?Generalist graph anomaly detection (GAD) aims to build a foundation model that detects anomalies on arbitrary unseen graphs without retraining or fine-tuning. Sufficient data are essential for foundation model training,…arXiv · MLLearning to Plan by Looking Back: Hindsight Hierarchies for Training Reasoning ModelsWe introduce a self-improvement loop for reasoning models based on the following observation: Even when the difficulty of a problem exceeds the model's current solving abilities, an additionally supplied solution might e…arXiv · MLA Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box OptimizationBlack-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining…arXiv · MLJust Weather Scoring: Efficient End-to-end Nowcasting with Distributional DiffusionGenerative diffusion models are well-suited for probabilistic precipitation nowcasting, but existing approaches often rely on separately trained compression or deterministic forecasting components and remain costly at in…arXiv · MLDataSense-Bench: The First Step Toward an AI ScientistAs claims about recursive self-improvement (RSI) and artificial general intelligence (AGI) proliferate, we ask a simple question: do frontier AI models have a sense of data, i.e., can they reliably select the right data…arXiv · MLQuickest Change Detection with Diffusion-Integrated ScoresClassical CUSUM relies on the log-likelihood ratio of the underlying distributions, which cannot generally be computed from finite pre- and post-change samples alone. We propose diffusion-integrated score CUSUM (DI-SCUSU…arXiv · MLSciTBERT: A family of chronologically consistent language models for scientific and technological language processingPre-trained transformer models are increasingly being used to study scientific and technological progress. Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression…arXiv · MLVerification with Transfer: Exact Information Frontiers and Their Price in CallsA verifier that accepts or rejects whole answers reveals little: under a flat prior over $k$-bit answers, zero error needs $2^k-1$ verifications. The usual remedy is to solve related source tasks, either all first, as a…arXiv · MLISBO: Scalable Spatio-Temporal Bayesian Optimization with Log Gaussian Cox Process Models via the INLA-SPDE ApproachBayesian Optimization (BO) is a popular method for efficiently optimizing expensive black-box objectives. However, BO utilizing standard Gaussian Processes is ill-suited for doubly stochastic Cox Processes that are often…arXiv · MLTraining on the Future: A Delay-Aware Audit of Test-Time Adaptation for Time-Series ForecastingTest-time adaptation (TTA) methods for time-series forecasting update a deployed model, or a small adapter around it, from incoming ground truth. But the label of an $H$-step forecast exists only $H$ steps later, and rea…arXiv · MLAdaCast: Conditional Parameter Generation for Adaptive Time Series ForecastingTime-series foundation models (TSFMs) have achieved strong forecasting performance across domains. However, most adaptation methods remain static. Existing all-in-one methods learn a single set of dataset-level parameter…arXiv · MLRIFT: Relative Isolation From Trees For Anomaly DetectionIsolation Forest (IF) is a widely used baseline for unsupervised anomaly detection. Recent studies provide a closed-form expression for the infinite-forest limit for one-dimensional data. Inspired by the geometric interp…arXiv · MLBatch Before You Lift: Scalable Topological Deep Learning on Large GraphsTopological Deep Learning extends graph-based learning to higher-order domains, such as hypergraphs, cellular, and simplicial complexes. These domains are typically constructed from patterns in an input graph through a p…arXiv · MLAdaptLSTM: Efficient Adaptive Online Learning for Cloud Workload Forecasting under Distribution DriftAccurate workload forecasting is critical for elastic resource provisioning in web-scale cloud services, where distribution shifts driven by viral content, product launches, and user behavior degrade offline-trained mode…arXiv · MLUnlocking the Regulatory Genome by ARGUS: An Evidence-Constrained Agentic Framework for Interpreting Single Nucleotide VariantsOver 90% of disease-associated variants from genome-wide association studies fall in noncoding regulatory regions, yet their functional interpretation remains a central open problem in genomic medicine. Large language mo…arXiv · MLSpatial Pattern Formation from Multi-Agent Learning in Public Goods DilemmasSpatial public goods models show that prescribed movement toward richer locations can generate spatial patterns. We ask how such patterns emerge when agents learn where to move and how learning rates shape their conseque…arXiv · MLPrior or Feedback? What an LLM Uses When Adapting Neural OperatorsDo LLM scientific agents rely only on their initial task context, or do they adapt their decisions in response to experimental feedback? We study this question in neural operator adaptation, where a large language model…arXiv · MLSparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM InferenceThe memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency. Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution t…arXiv · MLComposite Online-to-Nonconvex Conversion with Optimal Oracle ComplexityWe consider stochastic nonsmooth nonconvex composite optimization, which includes several important problems such as constrained optimization and the regularized training of neural networks. The objective is the sum of a…arXiv · MLPrediction-Powered Data Fusion for Treatment Effect EstimationRandomized controlled trials (RCTs) identify treatment effects without confounding but are often small, whereas observational studies (OBS) are large but may be confounded. Many estimators combining a small RCT with a la…arXiv · MLRiCo: Neural Simulation of Rigid-Body Interactions via Local Contact ReasoningAccurate simulation of rigid-body interactions is essential for predictive physical world models. Despite recent progress in modeling object dynamics, capturing how local contacts between surfaces shape object motion rem…arXiv · MLasdex: Automatic Sparse Differentiation in JAXMany tasks in scientific computing and machine learning require the Jacobian or Hessian matrix of a function. Automatic differentiation (AD) computes these derivatives to machine precision, but materializing a dense $m \…arXiv · MLVFold: Symmetry-Aware Cross-Layer Value Cache CompressionWhile caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths. One solution is to compress this memory by exploiting inter-layer cache s…arXiv · MLAmbient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient LearningWe introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications. RefineMix uses out-of-distribution data at selected diffusion times t…arXiv · MLOvercoming Prior Barriers: Supervised Fine-Tuning under Long-Tail DistributionSupervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support. Frequent concepts are more like…arXiv · MLSplitJEPA: Learning Invariant and Variant Latent Worlds without ReconstructionUnderstanding a dynamical world calls for more than a latent state that summarizes its observations: the state should also be organized into the factors that stay shared across related observations and the factors that v…arXiv · MLSubspace Uncertainty and Sharp Sampling Thresholds on the Boolean CubeWe study Gaussian regression under squared population $L_2$ loss in a known $m$-dimensional subspace of degree-at-most-$k$ functions on the $d$-dimensional Boolean cube. Random inputs can undersample regions essential fo…arXiv · MLClosing the Horizon Gap in Policy Optimization for Adversarial MDPsWe consider policy optimization for online episodic tabular Markov decision processes (MDPs) with adversarial losses and bandit feedback. Policy optimization updates the policy locally at each state and avoids optimizati…arXiv · MLBilevel optimization for data-driven learning of Koopman embeddings using kernel-based autoencodersKoopman operator theory provides a linear framework for analyzing nonlinear dynamical systems and has become a major tool for data-driven modeling. A central challenge, however, is that finite-dimensional approximations…arXiv · MLOnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal TransportAgents are deployed in applications from trip planners and stock trading to IT incident triage. In most cases, LLM agents work autonomously with minimal rule-based safeguarding, leading to cost and safety issues from irr…arXiv · MLMarformer: A Transformer for Predicting Missing Data DistributionsReal decisions are made under incomplete information. If we observe only some of the random variables we need, we can predict the others. The \textbf{conditional marginals} over the missing variables are the key ingredie…arXiv · MLLong Text to Predictive Features: LLM-Guided Blockwise Feature Engineering via Executable Program SearchIndustrial risk-control systems typically rely on structured-data models for efficient prediction, yet substantial valuable information remains embedded in unstructured long text. Extracting this information through manu…arXiv · MLHRIL: Learning Multimodal Synergy via Higher-Order Tensor ModelingSelf-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unli…arXiv · MLProspective Prediction of OOD Degradation from Source-Side Training DynamicsWe study whether persistent out-of-distribution (OOD) degradation can be predicted before it is directly observed using only source-side training dynamics. In a controlled shortcut-learning setting, a simple logistic reg…arXiv · MLLearning Kilometer-Scale Weather Prediction with Global-Regional AlignmentKilometer-scale regional weather forecasting is essential for local weather warnings and weather-sensitive decisions. Existing data-driven approaches often rely on numerical forecasts for large-scale guidance or require…arXiv · MLPredicting Alignment Generalization with Value RepresentationsLLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highl…arXiv · MLWOVEN: Weaving Visual World Modeling into Multimodal LLMsMultimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether thi…arXiv · MLA Unified Bellman Operator for Safety-Critical Reinforcement LearningReinforcement learning in safety-critical domains requires maximizing task performance while strictly adhering to safety constraints. Existing safe reinforcement learning paradigms typically force a trade-off: they eithe…arXiv · MLBeyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence MatchingDense correspondence matching has historically been bounded by simplifying spatio-temporal priors, such as smooth motion and rigid geometry. While effective for classical tasks, these assumptions break down in image edit…arXiv · MLToward Joint Optimization of Circuit Depth and Training Data Size in Adaptively Grown Quantum ClassifiersBuilding a quantum model involves a tradeoff: how complex the circuit should be, and how much training data it needs. Caro et al. show that models with fewer trainable gates need less training data to generalize well. Q-…arXiv · MLFAITH: Feasibility-Aware Safety-Filtered RL for High-Dimensional SystemsSafe reinforcement learning commonly places safety and task performance in the same policy objective, where they can introduce competing updates. Safety filters separate them at action execution, but classical designs re…arXiv · MLVioLA: Learning Generalist Humanoid Control Policies from Human DataTeaching a humanoid to follow instructions with its whole body runs into two obstacles. Its action space is large and tightly coupled: legs, arms, and fingers must move together while the robot keeps its balance, which m…arXiv · MLDensity Ratio Estimation with Stein Displacement FieldsDensity ratios quantify distribution shift from a probability-mass point of view, whereas displacement fields describe, from a dynamical point of view, how one distribution is transported onto another. Although both offe…arXiv · MLRounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State QuantizationQuantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates. We redesign 4-bit optimizer-state quantization for…arXiv · MLCaught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized DeceptionRecent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people. We show that white-box deception detection via probes can be scaled up to frontier monitoring settings b…arXiv · MLBi-FORK: Generative Modeling of High-Dimensional Bifurcating SystemsBifurcations are ubiquitous in physical systems, from structural buckling to fluid and climate dynamics, yet they remain largely unexplored in deep learning. At a symmetry-breaking bifurcation, a single input admits mult…arXiv · MLOne Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed ExpertsIn this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores de…arXiv · MLA Balanced Data Diet: Addressing the Exploration Bottleneck in Mega-Scale RL for Robot ControlGeneral-purpose robots must perform a wide range of tasks from agile locomotion to dexterous manipulation. While sim-to-real reinforcement learning (RL) has proven to be a useful tool for this goal, current RL pipelines…arXiv · MLCSF: Contextual Safety Filtering for Motion GeneratorsText-conditioned motion generators produce trackable whole-body motion, but they have no notion of scene-dependent safety: the same action may target an object or a person. Existing safeguards either inspect the prompt,…arXiv · AILearning to Plan by Looking Back: Hindsight Hierarchies for Training Reasoning ModelsWe introduce a self-improvement loop for reasoning models based on the following observation: Even when the difficulty of a problem exceeds the model's current solving abilities, an additionally supplied solution might e…arXiv · AIUnifying Policy Learning and State Prediction through Spatial Language ModelingLearning how actions change scene geometry can provide complementary supervision for goal-directed manipulation. We introduce Spatial Language Modeling, which represents scene contours, goals, action targets, and future…arXiv · AIRecursive Self-Improvement through Multi-Agent Self-SupervisionRecursive self-improvement (RSI) of a model on non-verifiable tasks, such as open-ended research, faces a supervision bottleneck when its outputs exceed what even human experts can reliably assess, leaving the model itse…arXiv · AIAI-Based On-Board Maritime Object Detection for Earth Observation Payload Data Reduction on Versal Embedded HardwareVery-high-resolution Earth-observation satellites acquire more data than they can store and downlink, while in maritime surveillance the vessels cover a tiny fraction of each scene. We study onboard vessel detection as a…arXiv · AIA Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box OptimizationBlack-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining…arXiv · AIWhen Has a Bayesian Neural Network Sampled Enough? Adaptive Inference Time with Statistical GuaranteesBayesian neural network predictions are commonly approximated using a fixed number of Monte Carlo samples per input, without controlling the resulting error that comes from this finite sample. We propose the use of confi…arXiv · AIVibeEdit: Image Editing with Canvas InstructionsIn text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike. We introduce a new image…arXiv · AIFrom Prompting to Composing: A Spatial Canvas Interface for Poster GenerationText prompting is an indirect interface for poster generation, requiring users to encode inherently two-dimensional composition intent into a one-dimensional sequence of words. We introduce a Spatial Canvas Interface tha…arXiv · AIReSI: Recursive Safety Improvement toward Resistant and Resilient AIRecursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment.…arXiv · AILa-Ribo: RNA Co-Design via Geometry-Latent Flow MatchingRNA function arises from the coupling of nucleotide sequence and three-dimensional structure, motivating their joint design. Coordinating global folding with nucleotide-level detail remains challenging under limited stru…arXiv · AIAdaCast: Conditional Parameter Generation for Adaptive Time Series ForecastingTime-series foundation models (TSFMs) have achieved strong forecasting performance across domains. However, most adaptation methods remain static. Existing all-in-one methods learn a single set of dataset-level parameter…arXiv · AIBatch Before You Lift: Scalable Topological Deep Learning on Large GraphsTopological Deep Learning extends graph-based learning to higher-order domains, such as hypergraphs, cellular, and simplicial complexes. These domains are typically constructed from patterns in an input graph through a p…arXiv · AIReal-Time Motion Planning with Dynamic Hazards: Classical vs. Learning-Based MethodsWe study real-time motion planning in dynamic hazard fields through a controlled comparison between classical planning and learning-based methods. Rather than introducing a new planner, we construct a unified benchmark i…arXiv · AIHow Much Audio Is Left In An Embedding? An Inversion Audit Of Audio EncodersPretrained audio encoders are reused for downstream tasks that are often unknown when the encoder is trained, so their usefulness depends partly on which signal properties survive the pretext objective. We study this ret…arXiv · AISyn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal EmbeddingsOmnimodal embeddings naturally involve both shared representations and modality-specific features across heterogeneous inputs. However, existing omnimodal embedding methods often rely on a single shared parameter space o…arXiv · AIDVD: Dynamic Vector Decoding for Efficient MLLM-based PerceptionMultimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving. However, ex…arXiv · AIUnlocking the Regulatory Genome by ARGUS: An Evidence-Constrained Agentic Framework for Interpreting Single Nucleotide VariantsOver 90% of disease-associated variants from genome-wide association studies fall in noncoding regulatory regions, yet their functional interpretation remains a central open problem in genomic medicine. Large language mo…arXiv · AIOne Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent GuardrailsA typed decision model reads a piece of text and returns a probability over caller-defined options, each with a short written definition, generating no text. Recent work places these models in agent systems as guardrails…arXiv · AIMachine Learning Meets High-Energy Nuclear Physics: From Pattern Recognition to Physics-Integrated DiscoveryMachine learning (ML) in high-energy nuclear physics (HENP) is entering a new stage in which physical knowledge is incorporated more directly into data analysis, simulation, and physics inference. This mini-review focuse…arXiv · AIMulti-Agent Egocentric World Model with Fine-Grained Embodied InteractionEgocentric world models predict first-person observations conditioned on an agent's actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interact within a shared en…arXiv · AILearning Probabilistic Logic Programs with Functional Gradient Guided Language ModelsDeclarative logic programs offer a powerful and interpretable abstraction for encoding relational structure and neurosymbolic reasoning, by expressing dependencies as weighted compositional rules. However, inducing them…arXiv · AILooking Inside LLMs: Small-World Connectivity as a Signature of Reasoning PerformanceUnderstanding large language model (LLM) reasoning requires looking beyond behavioral performance to examine how reasoning ability is reflected in internal organization. Inspired by neuroscience findings linking higher i…arXiv · AIA Structural Theory of Cognitive Representation and Problem Solving,Contexts, Invariance, and the Knowledge SpaceLearning and problem solving depend critically on the structure of internal representations. While many modern data-driven artificial systems achieve strong predictive performance, their learned representations often lac…arXiv · AIVerdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance SystemsLarge language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given. We test this directly across five models and 20 regulatory and platform-policy domains: de…arXiv · AIPrior or Feedback? What an LLM Uses When Adapting Neural OperatorsDo LLM scientific agents rely only on their initial task context, or do they adapt their decisions in response to experimental feedback? We study this question in neural operator adaptation, where a large language model…arXiv · AIPrediction-Powered Data Fusion for Treatment Effect EstimationRandomized controlled trials (RCTs) identify treatment effects without confounding but are often small, whereas observational studies (OBS) are large but may be confounded. Many estimators combining a small RCT with a la…arXiv · AIRiCo: Neural Simulation of Rigid-Body Interactions via Local Contact ReasoningAccurate simulation of rigid-body interactions is essential for predictive physical world models. Despite recent progress in modeling object dynamics, capturing how local contacts between surfaces shape object motion rem…arXiv · AIContiLNN: Mitigating Slice Sampling Discontinuity with Liquid Neural Networks for Medical Image RestorationAnatomical continuity provides complementary information for medical image restoration, but its use requires accounting for local anatomy and variations in slice sampling. We introduce ContiLNN, which augments two-dimens…arXiv · AICan AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent CompetitionAdversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging. AI agents offer an alternative by turning game experie…arXiv · AIOvercoming Prior Barriers: Supervised Fine-Tuning under Long-Tail DistributionSupervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support. Frequent concepts are more like…arXiv · AIDistilling Routed 3D Privilege for Spatial Reasoning in Vision-Language ModelsSpatial reasoning remains a persistent weakness of vision-language models (VLMs), because RGB inputs do not directly provide geometric evidence. Existing remedies either inject 3D into the model at inference, paying arch…arXiv · AIAccurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge ConflictWhen retrieved evidence contradicts an agent's prior beliefs, does it revise its answer, acknowledge uncertainty, or persist with an incorrect conclusion? Existing evaluations of agentic systems focus primarily on task s…arXiv · AICited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought FaithfulnessLarge language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we s…arXiv · AIHANS: A Handwritten Answer Sheet Dataset for Noisy Hybrid Document ParsingIntelligent grading and automated scoring technologies constitute critical infrastructure for smart education. However, existing document parsing and handwriting recognition benchmarks are predominantly designed for well…arXiv · AIOnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal TransportAgents are deployed in applications from trip planners and stock trading to IT incident triage. In most cases, LLM agents work autonomously with minimal rule-based safeguarding, leading to cost and safety issues from irr…arXiv · AIARC: A Reasoning Recipe for Robot Foundation ModelsThe prevailing approach to improving robot foundation models (RFMs) relies on larger models, more robot demonstrations, and costly training at scale. We show that there exists an effective and efficient complementary app…arXiv · AIGeoReform: Reflective Formalization Evolution for Multimodal Geometry Problem SolvingMultimodal large language models (MLLMs) often struggle to identify and use geometric relations in diagrams. Recent methods address this challenge by converting geometric entities, relations, and constraints into explici…arXiv · AIHRIL: Learning Multimodal Synergy via Higher-Order Tensor ModelingSelf-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unli…arXiv · AISpaceFlow: Locally Controllable 3D GenerationCurrent 3D generation methods lack explicit local control: geometric adherence is often defined by a global control strength, and appearance cannot be specified locally. We present SpaceFlow, a training-free pipeline for…arXiv · AILeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPCWorld action models (WAMs) predict actions and future observations, typically from a reconstruction-based representation that carries noisy, redundant information which can complicate downstream predictions. We introduce…arXiv · AISearching for "Harmful Refusal": A Psychometric Audit of an AI Safety BenchmarkSafety benchmarks typically report one overall score for a suite of datasets, each of which may target one or more safety-related attributes, so models with similar overall scores can have very different attribute profil…arXiv · AIPredicting Alignment Generalization with Value RepresentationsLLM developers post-train their models to exhibit prosocial values and behavioral traits, which are enumerated in an alignment target. However, while recent post-training developments have yielded models that score highl…arXiv · AIMAMHOI: Factorizing Scene-Aware Human-Object Interaction through AffordancesGenerating realistic human-object interactions (HOI) in complex 3D scenes requires two complementary capabilities: reasoning about interaction feasibility in the environment and synthesizing realistic human-object motion…arXiv · AIRoboRSI: Stable, efficient, and reusable robot self-evolution in complex real-world environmentsA generalist robot should not only perform diverse tasks but also improve through experience, turning what it learns during execution into capabilities that later tasks can reuse. Robot agents that act through code can a…arXiv · AIEcology of AI Agents: Collaboration Creates a Population Threshold for TakeoffAI agents can now conduct real-world cyberattacks, scale up capabilities with the number of agents, and collectively pursue misaligned goals to obtain rewards. Together, these factors raise the risk of a population explo…arXiv · AICaught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized DeceptionRecent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people. We show that white-box deception detection via probes can be scaled up to frontier monitoring settings b…arXiv · AIBi-FORK: Generative Modeling of High-Dimensional Bifurcating SystemsBifurcations are ubiquitous in physical systems, from structural buckling to fluid and climate dynamics, yet they remain largely unexplored in deep learning. At a symmetry-breaking bifurcation, a single input admits mult…arXiv · AIBrickBench: Evaluating Agentic Brick DesignWe propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be ph…arXiv · AIFrom Reactive Containment to Proactive Assurance: Lessons from OpenAI, Anthropic, and Google Agent Security IncidentsIn 2026, cybersecurity evaluations involving OpenAI, Anthropic, and Google agents reached real systems outside their authorized test scope. The paths were different. OpenAI agents exploited research infrastructure, coord…arXiv · AIOn the estimation and validity of AI time horizons---a statistical look at the METR plotMETR's 50\% time horizon measures the human completion time of software tasks that an AI solves with 50\% probability, allowing AI capabilities to be expressed in interpretable units. On 228 tasks and 26 AIs, we recomput…AppleNormalizing Trajectory ModelsDiffusion-based models decompose sampling into many small Gaussian denoising steps, an assumption that breaks down when generation is compressed to a few coarse transitions. Existing few-step methods address this through…AnthropicBuilding on our commitment to American scientific discoveryAnthropic2026 Usage Policy updateAnthropicIntroducing the Anthropic Cyber MissionOpenAIGPT-6 and Intelligent UI for everyoneGPT‑6 is rolling out globally in ChatGPT with Intelligent UI, delivering faster responses with visuals and interactive experiences you can explore and use directly.OpenAIRadisson Hotel Group brings hotel discovery into ChatGPTRadisson partnered with Accenture to build a ChatGPT plugin using OpenAI technology, helping travelers find, compare, and book hotels while planning their trips.OpenAIHelping teens learn, plan, and shape the future of AICollege Planner is coming to ChatGPT for Teens to help students manage college applications, alongside new flashcards, quizzes, and a teen AI council.NVIDIAFaster Scientific Image Analysis with NVIDIA cuPhotonObservatories and telescopes, lasers and X-ray light sources, and other high-throughput instruments generate image data faster than CPU-bound pipelines can...NVIDIAScaling Decision Optimization to 100 Million Variables and Beyond with mPDLP in NVIDIA cuOptSupply chain problems are expanding across more SKUs, lanes, and constraints than ever before, while energy grids are balancing more distributed sources in real...NVIDIAValidate AI Factory Changes with Digital Twins and AI AgentsAI factories are some of the most complex operations in the world, combining GPUs, CPUs, switches, DPUs, and SuperNICs alongside schedulers, orchestration...NVIDIAThe Machines that Make the MachinesHow we taught robots to assemble GB300 tester trays and what it taught us about robot learning, mechanical intelligence, and good old-fashioned engineering The...Microsoft ResearchAgent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real HarnessesTraining AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it ea…MicrosoftFrom Spec-First to Enterprise-Ready: Extending GitHub Spec KitHow presets, extensions, bundles, workflows, and governed catalogs make Spec-Driven Development adaptable at scale This article continues Spec-Driven Development: A Spec-First Approach to AI-Native Engineering, which int…Hugging FaceOne Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMOHugging FaceIntroducing Falcon ASRHugging FaceMultimodal open d1 decision models for the edgeGitHubSecret protection must scale with softwareDevelopers aren’t becoming more careless; they’re being outpaced. The tools that let developers create more software should also take on more of the work of protecting it. The post Secret protection must scale with softw…CloudflareBuilding an evidence-grounded agentic security operations harness on CloudflareCloudflare Managed Defense uses a team of specialized AI agents built on Workers and global network telemetry to analyze security alerts. By separating deterministic evidence collection from model inference, the system d…PinterestMetrics Board: Building an Agent-ready Metrics LayerMichele Ceccacci; Software Engineer I | Jason Coffman; Sr. Software Engineer | Colm O’Shaughnessy; Software Engineer II | Laura Palmer; Staff Product Manager | Adam Podraza; Manager, Engineering | Surya Karri; Manager, E…OpenAIAtlassian and OpenAI expand partnership to turn enterprise knowledge into actionAtlassian and OpenAI are expanding their partnership to connect frontier models with enterprise knowledge and help teams plan, build, and deliver work.OpenAIAdvancing computer use with IroncladLearn how OpenAI and Ironclad are training and evaluating AI agents on complex contracting workflows to advance computer use for professional work.OpenAISharing AI progress in mathematicsOpenAI publishes new results on open problems in mathematics from an internal frontier model and shares Lean proof formalizations and research details on GitHub.OpenAIHow Jump Trading is scaling quant research with ChatGPTJump Trading uses OpenAI to expand quantitative research. See how longer-running AI workflows combine multiple data sources with human review.NVIDIAControl How Your GPU Shares Work with Green ContextsGPU applications increasingly consist of multiple independent components running at the same time within a single process: a latency-sensitive operator...NVIDIAAICR v1.0: Open, stable, and verifiable GPU cluster configurationGPU-accelerated Kubernetes clusters depend on compatible versions across dozens of components, each on its own release cycle: host kernels, GPU drivers,...NVIDIAHow DOCA GPUNetIO Unifies GPU-Initiated Networking Across the NVIDIA Software StackGPU applications increasingly need networking and data movement to behave like first-class GPU-controlled operations rather than host-driven services. When the...NVIDIAScale Bitwise-Deterministic Pretraining with NVIDIA Megatron CoreBitwise determinism makes large-scale pretraining easier to debug, validate, and resume reproducibly. These benefits become especially valuable when training...Microsoft ResearchWhat AI gets wrong and what failure teaches usJennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures”…Mistral AIIntroducing Mistral Large 4MicrosoftWhat is Agent Experience (AX)?Your agent says it's done, and the code compiles, but did it pick your technology? Did it use it correctly? Learn what Agent Experience (AX) is, how to measure it, and why the obvious fix isn't always the right one. The…MicrosoftBuild Azure canvases with the Canvas authoring pluginLearn how the Canvas authoring plugin helps you build Azure canvases with Microsoft Canvas Toolkit, from a read-only starter to a working app you can inspect, test, and adapt. The post Build Azure canvases with the Canva…MetaNTS: Authenticated Time at MetaMeta’s public time service now speaks NTS (Network Time Security, RFC 8915) at nts.meta.com. Packets are authenticated, so a device can verify the time came from us and was not modified on the way. Our NTS servers hold n…Hugging FaceFalcon-Emirati: When an LLM Learns the Dialect, the Culture, and the NuanceGrabSmarter personalization: How property data helps us understand user price sensitivityIntroduction Existing systems estimate user price sensitivity primarily from spending behavior or demographic proxies. They do not systematically account for residential property values, which can indicate a user’s finan…GitLabEvery artifact your teams ship, assembled right the first timeEvery software build is assembled from open source packages, base images, and libraries, with your code on top. When one of those components is missing or has drifted, the build fails. At agent speed, those failures mult…GitLabDependency Firewall: Block risky packages before the buildAttackers disguise malicious packages as ones you trust. In June 2026, GitLab researchers found five malicious PyPI packages , four of them typosquats of Flask, Requests, and NumPy, that run at install time and steal CI/…GitLabGitLab Transcend: Speed you can trust, all the way to productionAt Transcend, we made over a dozen announcements across all four layers of our architecture for agentic software engineering: agent orchestration, data and context, DevOps workflows, and governance and security. Watch th…GitHubBuilding Git infrastructure for agent-scale developmentWe're rebuilding GitHub's Git infrastructure while GitHub keeps running, creating a foundation for agent-scale software development. The post Building Git infrastructure for agent-scale development appeared first on The…DiscordDiscord Patch Notes: October 6, 2026Check out the finer details of the more technical fixes implemented into Discord recently.Google DeepMindEmbeddingGemma 2: an open, lightweight multimodal embedding modelCohereMultilingual Bridges: How Data Mixing Unlocks In-Language ReasoningCloudflareThe keys to the Internet change on October 11. Are you ready?On October 11, 2026, the DNS root switches to a new key-signing key (KSK-2024). Learn what this means for you, and how RFC 8509 trust anchor sentinels allow you to test whether your DNS resolver is ready for the rollover…AWSAll the numbers: Amazon Prime Day 2026 powered by AWSAmazon Prime Day 2026 was exclusively for Prime members and ran June 23–26, 2026 with millions of deals across more than 35 categories. As part of our annual tradition of sharing how AWS powered Prime Day (2016, 2017, 20…AppleRISED: Rubrics for Agentic Multi-Environment Selection and Self-DistillationTraining a single LLM agent jointly across diverse interactive environments has attracted increasing attention as a route to generalist agents. Existing curriculum and data-selection strategies often allocate training at…AnthropicExpanding the Cyber Verification ProgramAirbnbBeyond synthetic testing: Capturing and replaying real database workloads at AirbnbHow we capture real production database traffic at Airbnb and replay it offline to load-test, plan capacity, and de-risk upgrades. By: Zuofei Wang , Erluo Li Introduction At Airbnb, MySQL-compatible databases are a criti…SalesforceAI Agents Need Memory: Building Persistent Memory for AgentforceBy Roopang Chauhan, Sundar Vedula, Peng-Wen Chen, and Nikhil Bojja. In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Roopang Chauhan,…OpenAIBuilding advertising for the way people use AIOpenAI introduces a new visual ad format in ChatGPT and expands measurement tools, attribution partnerships, and brand suitability for advertisers.OpenAIOur approach to EU text provenance rulesHow OpenAI is approaching text watermarking under EU rules. Learn where watermarks apply, how detection works, and why access starts with researchers.GrabPowering AI-led research through simulationThe short story Consider a Friday evening. A food order arrives from a mall in the city center. One driver is nearby; another is finishing a drop-off and will be available shortly; a second order from the same mall may o…GitLabTwo front doors: Module-level access in a Django GRC appGitLab s engineering team builds a lot of our own internal tooling, but one platform in particular forced us to rethink how we handle authorization: our internal GRC tool that serves two very different customer bases und…GitHubReviewBench: An open benchmark for AI code reviewWe’re launching ReviewBench, a benchmark for code review agents built on representative GitHub pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics. The post ReviewBench: An ope…CohereCohere and PwC partnership announcementCohereNorth 2: Enterprise AI Without CompromisesCloudflareOne year later: the power of 1.1.1.1 internsA year after announcing our goal to hire 1,111 interns, more than 750 early-career builders have shipped real products across 48 teams at Cloudflare. From Birthday Week launches to post-quantum security, our interns prov…CloudflareEverything we launched during Birthday Week 2026We celebrated our 16th birthday with 46 announcements across open source, post-quantum security, AI agents, and developer platform upgrades. Here’s a day-by-day roundup of everything we shipped.AWSAWS Weekly Roundup: Amazon Bedrock Managed Agents powered by OpenAI, Q3 service availability updates, Kiro workflows, and more (October 5, 2026)Last week, we announced a public preview of Amazon Bedrock Managed Agents powered by OpenAI, built on a customized version of OpenAI’s Agents API engineered to be AWS-native and integrated with AWS resources. You can now…AppleNegotiating Ontological Boundaries in User-Authored Personal Sensing SystemsDesigned artifacts are ontological, shaping, and at times limiting, what becomes possible or imaginable. One path toward mitigating such foreclosures is giving people power over how systems are designed and built. Despit…Hugging FaceThe Agent Said It Was Done. The Database Disagreed.OpenAIChatham scales its capital markets expertise with OpenAIChatham Financial uses Codex and GPT-5.6 to build technology and redesign workflows, cutting trade validation from 30 minutes to under 4.OpenAIA model guide for the GPT-6 familyLearn how startups can choose GPT-6 models, tune reasoning effort, improve prompts and skills, coordinate tools, and prepare workflows for production.Hugging FaceAutoSynthData: Generating Training Data for Enterprise AgentsHugging FaceOpen-sourcing AstaBrief, the fast report-generation model in AstaGitLabDeepSeek-Reasonix: How a poisoned config can hijack an AI coding agentGitLab s Threat Research Group discovered a command execution vulnerability ( GHSA-grg2-7gc6-36m6 , CVE-2026-102437 ) in DeepSeek-Reasonix Studio, a desktop git client designed for developers pairing with AI coding assis…GitHubAI is changing developer work. Here are three skills to strengthen.Learn to direct AI agents, critically review their output, and keep technical judgment at the center of your workflow. The post AI is changing developer work. Here are three skills to strengthen. appeared first on The Gi…Cloudflare2026 Birthday week: network performance updateCloudflare now ranks as the fastest provider across 74% of the top 1,000 global networks. By incorporating background telemetry from Cloudflare Challenge Pages, we have expanded our real-user measurement scale while main…CloudflareBuilding for good: How civil society organizations are automating on CloudflareSome of the world's leading organizations are building the future of non-profit work with Cloudflare.CloudflareProtected Quick Tunnels: simple accountless authentication for your next dev projectQuick Tunnels now support email authentication. Add --allowed-mail to one cloudflared command, and only the addresses or domains you list can reach your local app. No Cloudflare account required on either side.CloudflareFollow the thread: a new dashboard to investigate account abuseFraudsters are increasingly using AI to bypass stateless security checks. Cloudflare's new Account Abuse Protection dashboard uses stateful analysis and edge-generated Hashed User IDs to help teams investigate and block…CloudflareAnnouncing Cloudflare OHTTP Gateway – expanding access to Cloudflare’s privacy-preserving infrastructureWe’re announcing the closed beta of a self-serve Cloudflare OHTTP Gateway. We’re also renaming our Privacy Gateway to Cloudflare OHTTP Relay to better distinguish the two products.CloudflareUpdates on our pledge to make Cloudflare features accessible to everyoneA year after pledging to eliminate two-tier product access, Cloudflare has expanded Logpush, multi-account governance, and higher platform limits to all accounts. Here is an update on our progress, how we dogfood these t…CloudflareIntroducing Cloudflare Traces: follow requests through our entire platformCloudflare Traces shows how a request moves through security rules, transformations, cache, routing, Workers, and your origin, then follows it across services running anywhere in your stack.CloudflareIntroducing Web Search API via AI GatewayCloudflare AI Gateway now supports native web search API integration in partnership with Ceramic.ai, Exa, and Linkup. Developers can now inject real-time web context into model inference calls via AI Gateway, REST APIs,…CloudflareStreamline: custom video pipelines with Cloudflare Stream and WorkersStreamline demonstrates how to build long-running, continuous video processing pipelines by pairing Cloudflare Workers and Durable Objects with a containerized media engine.Cloudflare8 major updates to Cloudflare ObservabilityCloudflare is launching eight major updates that bring logs, traces, analytics, alerts, dashboards, querying, and telemetry export into one observability platform, with simpler and more predictable pricing.AppleLimits of Confidence in DiffusionDiscrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step, drawing each from a per-position distribution and choosing which positions to write fr…AppleLanguage Discrimination Improves Linguistic Learning in Multilingual Speech ModelsMultilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening…AnthropicClaude Frontier Academy: $100M to train 10,000 engineersStripeWhy I tried to kill token billing (and why we kept it)Token billing is useful infrastructure, but usually a bad customer-facing pricing model. Your invoice should define the value your product delivers, not break down what it cost you to create it.OpenAIThe Den frees up 10-15 hours a week to grow with ChatGPT WorkAs it opens a new location, the social club prepares grant applications in 2 hours instead of 3 days and liquor-license materials in 3 hours instead of 4 days.OpenAIHow Albertsons Companies is reimagining retail from the inside outAlbertsons Cos. is using ChatGPT Enterprise and the OpenAI API to help teams work faster and make grocery shopping easier for millions of customers.OpenAIThe eternal complementAdvanced AI may matter most for the routine work behind breakthrough ideas. Explore why execution could shape the next economy and the pace of progress.NVIDIAFine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other LanguagesAutomatic speech recognition must handle how people actually speak, not only the languages and styles that dominate pretraining data. Regional dialects and...NVIDIABuild Local AI Apps with C++ and NVIDIA TensorRT RTX SamplesAdding AI models to local applications requires a portable model format, a reliable runtime, and acceleration that works across target systems. Do Inference Now...NVIDIABuild Applications on NVIDIA BlueField Faster with NVIDIA DOCA Agent SkillsAI agents are becoming a standard part of development workflows, but general-purpose agents weren't built with specialized infrastructure software such as...Hugging FaceIntroducing Olmo-core 3: Open, scalable training infrastructure for large MoEsGitHub10 technical talks I’m excited about at GitHub Universe 2026From verifying AI-written code to securing npm dependencies, these are the sessions I’m building my Universe agenda around. The post 10 technical talks I’m excited about at GitHub Universe 2026 appeared first on The GitH…DiscordCauldron Bubble w/o the Trouble? It’s Discord Halloween, WitchesA cold wind whistles, a chill sets in. It’s Discord Halloween: Cauldron Chaos edition! We’ve got frights, we’ve got delights, and we’ve even got some spooky surprises up our sleeve…CloudflareOne year later: Sovereign AI and the fight for choiceAI sovereignty is not a zero-sum game, but many governments now believe it is. Cloudflare's answer: more local open-source models, model-agnostic security tools, and a commitment to giving nations genuine choice.CloudflareIntroducing Clef: our open-source decision models, and new RL fine-tuning platformWe are introducing Clef and Clef-flash, open-source decision models hosted on Workers AI for high-speed classification and agentic workflows. Also launching: a new reinforcement learning platform that allows developers t…CloudflareSupport for modern cryptographic algorithms in WorkersCloudflare Workers is adding opt-in support for post-quantum-resistant algorithms ML-KEM and ML-DSA. You can try it today.CloudflareIntroducing Cloudflare Basin: an open, serverless data platform, now generally availableCloudflare Basin is built on Apache Iceberg and R2 Object Storage. With general availability, developers can ingest, manage, and query large datasets at scale without paying data egress fees.CloudflareWe want you to build the next Git platform on CloudflareCloudflare is hosting a competition to see who will build the next Git platform for an era of AI agents. Artifacts is in open beta, with Workers bindings, data jurisdiction controls, and event subscriptions for repositor…CloudflareAI Search is now generally availableAI Search is now generally available. It embeds image pixels directly for visual search, runs optical character recognition on scanned PDFs, accepts files up to 10 MiB, and works with any chat model. Here's what's new an…CloudflareIntroducing Workers KV Instant — powered by QuicksilverWorkers KV Instant delivers sub-2ms p99 read latencies and 250ms global replication across Cloudflare’s 300+ edge locations. KV Instant eliminates cold-read penalties and uses the familiar Workers KV API.CloudflareCloudflare OS: your company’s agent workspace, managed for youCloudflare OS gives everyone in your organization an agent workspace that knows how your company works and connects to its data and systems. We’re opening the waitlist for fully managed deployments that you’ll be able to…CloudflareAnnouncing Cloudflare K2: serverless event streamsCloudflare K2 is a serverless event streaming service built directly on top of R2 object storage for high-scale data movement and long-term retention. By decoupling producers and consumers at the edge, K2 enables durable…AWSAnnouncing AWS Well-Architected Agent, an AI-powered intelligence to optimize your cloud environment (preview)Introducing the public preview of AWS Well-Architected Agent, an AI-powered service that analyzes your AWS environment to deliver targeted, contextual recommendations for improving your applications' cost, security, perf…AppleRLTL;DR: Self-Improvement by Internalizing Self-Generated FeedbackThe common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-imp…AppleHow Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) pr…AnthropicBarclays scales Claude to upgrade operations and improve client experienceStripeOUSD is now the default stablecoin on StripeOpen USD (OUSD), a stablecoin built for global money movement, is now available across Stripe. Businesses can use OUSD to manage funds, make payments, and offer new financial services.SalesforceHow Deterministic Controls Turn AI Output Into Reliable Prompt TemplatesBy Vaibhav Raizada, Kumar Kasimala, and Ashish Gite. In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Vaibhav Raizada, a senior softwa…OpenAIHelping small businesses put AI to workOpenAI is partnering with America’s SBDC to expand hands-on AI training and local support for small businesses, alongside a new report on how small teams are using AI.OpenAIDisrupting a coordinated model-distillation campaignLearn how OpenAI disrupted a campaign to extract protected model reasoning and is strengthening defenses against adversarial distillation.NVIDIATracing Agent Harness Behavior with NVIDIA NeMo RelayAn agent can finish a task and still take an inefficient path. A failed search can trigger another search. A truncated file read can lead to a command fetching...NVIDIAExpanding AI Storage Access with NVIDIA cuObject and the NVIDIA SCADA Server SDKAI infrastructure engineers, storage developers, and cloud service providers need fast and secure access to high-capacity file and object storage to support AI...NVIDIADeploying an HSTU Generative Recommender with NVIDIA Dynamo-TritonGenerative recommender (GR) systems are emerging as a powerful new approach for large-scale personalization. Instead of treating recommendation as a set of...Microsoft ResearchForecasting space weather risks on power gridsExtreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arr…Hugging FaceOpen TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice CloningGoogle DeepMindIntroducing SynthID BioProof of concept for watermarking AI-generated proteins while preserving biological function.Google DeepMindGemini 4 Argon: our next era of frontier intelligenceCohereIntroducing Embed 5—A new family of frontier embedding modelsCohereRCP-nDCG@10: A New Approach to Enterprise Retrieval QualityCloudflareCloudflare Impact reaches $100 million in donationsThis week, Cloudflare's Impact programs will reach $100 million in donated services. This milestone means that thousands of entities including journalists, civil society, state and local governments, election management…CloudflareThe Internet has a second audienceMore than half the traffic reaching sites on Cloudflare is now automated, and AI agents are the fastest-growing part of it. We're giving site owners the tools to see who's visiting, decide who gets in, and charge for acc…CloudflareCloudflare Containers, rebuilt to scale agent sandboxesCloudflare Containers now start 6x faster, let your agent choose each sandbox's image and instance type at runtime, and support filesystem snapshots in public beta, all controlled from a Durable Object.CloudflareIdentify AI model overuse with User InsightsAI Gateway User Insights now adds task, model, turn, and user categories to help teams understand AI adoption and make better model decisions. This is available free to AI Gateway users.CloudflarePay Per Use: when AI uses your work, you should get paidPay Per Use is now in beta. AI companies report when they use publishers’ content, and Cloudflare handles billing, payouts, and reporting, so our customers are paid according to use.CloudflareMonetization Gateway beta: charge AI agents for consumption with HTTP 402Cloudflare’s AI Gateway, Ceramic.ai, Stocktwits, and more are using the Cloudflare Monetization Gateway today to charge agents for access to tokens, APIs, and MCP tools. U.S.-based sellers can now apply for access to the…CloudflareSimplifying domains for people and agentsCloudflare Registrar’s new search delivers fast, transparent results across 420+ extensions using Workers, Durable Objects, and WebSockets. Its expanded API and cf CLI also let agents search, register, and transfer domai…CloudflareDetect and send production issues straight to your agentYou can now use built-in error monitoring in Cloudflare Workers to group production failures and send stack traces, logs, traces, and application context directly to a coding agent to investigate further and open a pull…CloudflareCut your AI spend with AI Gateway's Auto RouterCloudflare AI Gateway now features a model router that evaluates request complexity using an edge-deployed classifier to select the optimal model. By balancing expected output quality against token costs, organizations c…AWSCelebrating Our Newest AWS Heroes – September 2026Today, we’re excited to introduce the newest members of the AWS Heroes program. AWS Heroes are a vibrant, worldwide group of AWS experts who go above and beyond to share knowledge, mentor others, and build thriving commu…AWSAmazon Aurora PostgreSQL now supports direct querying of Apache Iceberg and Parquet data in your data lakeAmazon Aurora PostgreSQL now lets you directly query Apache Iceberg and Parquet data stored in your data lake alongside live operational data—no ETL pipelines required. Powered by DuckDB embedded within Aurora, this capa…AWSAmazon S3 Vectors now supports metadata pre-filtering for higher recall on filtered searchesAmazon S3 Vectors now supports metadata pre filtering, delivering up to 5x higher recall on filtered searches. Filters evaluate before similarity search so scoped queries return more relevant results. Ideal for RAG, agen…AWSAmazon S3 Tables now support all Apache Iceberg V3 data typesAmazon S3 Tables now support all data types in the Apache Iceberg V3 spec, along with column default values, deletion vectors, and row lineage. Upgrade existing V2 tables or create new V3 tables with built-in compaction,…AppleSCLATE: A Substrate for Continual-Learning Agent Training and EvaluationContinual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and…AppleOn the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic StudyControlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive. Current approaches to conditioning are of…YelpThe making of the Yelp Assistant UIAs simple as it may look, a chat application is made up of thousands of small decisions that we make along the way, all with the goal to build a delightful experience for users. But each decision comes with a cost, often…StripeHelping personal agents shop more intelligently and reliably with LinkAs agents take on more purchases, agent builders have increasingly asked us to help agents navigate checkout and earn consumer trust. Today, we’re sharing three major improvements to support this broader set of needs.OpenAIIntroducing dotsDots by OpenAI are proactive assistants that can keep working across complex projects and everyday tasks. Learn how dots help you stay in control while work moves forward.OpenAIIntroducing GPT-6.1 SolMeet GPT-6.1 Sol: near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra’s standard API input and output token prices.OpenAIDevDay 2026 RecapExplore more than 20 announcements from OpenAI DevDay 2026, including GPT-6 Astra, ChatGPT, Codex, APIs, security, and new tools for builders.NVIDIALower the Cost of Building and Running Visual AI Agents with NVIDIA VSS Blueprint 3.3Vision-language models have made it possible to build visual AI agents that understand video at production scale. The harder problem is turning that capability...NVIDIAAI Native by Design: Lessons Learned from Building NVIDIA TensorRT Model ConnectParallel work, model-family isolation, reversible changes, and GPU-backed validation shaped an open source project designed around coding agents NVIDIA TensorRT...Microsoft ResearchIntroducing Quine: An AI research system designed for the complexity of biologyBiology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and…MicrosoftEnabling Consistent AI-Assisted Engineering with GitHub Copilot PluginsWhen an engineer or team creates useful agentic artifacts for their daily work (e.g., instructions, skills, and MCP configurations), others naturally want to adopt them when they see how those artifacts could help with t…MicrosoftBuild with Azure Canvases: a shared workspace for agentsSee how Azure Canvases bring plans, code, previews, and deployment feedback together in a collaborative workspace for agent-assisted development. The post Build with Azure Canvases: a shared workspace for agents appeared…Hugging FaceGetting the Source Right, Not Just the Fact: Source-Aware Verification for MCP AgentsHugging FaceNVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular PredictionGitLabGitLab and Claude Code: Fast, compliant AIGovernment agencies are feeling twin pressures: While the U.S. Office of Management and Budget (OMB) is urging you to deploy AI faster, the U.S. Government Accountability Office (GAO) wants guardrails in place before tha…GitHubDeveloper policy update: Transparency, state policy, and what’s aheadExplore GitHub’s latest transparency data and learn more about policy updates affecting developers and open source. The post Developer policy update: Transparency, state policy, and what’s ahead appeared first on The Git…EtsyThat Which We CAN Do: Sustainability in a World of AIAs computing use grows, so does its environmental footprint. While hardware and software have become significantly more efficient, AI demand is growing even faster, outpacing those efficiency gains and increasing the ove…DropboxEvolving our calendar assistant Reclaim to be AI-native without starting overHow we redesigned Reclaim for AI and natural-language requests while preserving the scheduling experience people rely on.CloudflareBuilding a certificate authority for the whole InternetTwelve years after launching Universal SSL, Cloudflare is applying to become a certificate authority. By combining an established root, an ACME-first approach, and Merkle Tree Certificates, we are building a post-quantum…CloudflareUsing AI to chart a course for our post-quantum migrationWe’re building CryptoLabe, an internal AI-powered tool that discovers cryptography across our codebase, surfaces dependencies, and helps us progress toward a full post-quantum migration by 2029. Here’s what we’ve learned…CloudflareAdaptive application security for the AI era: how Cloudflare connects code, traffic, and intelligence to stop attacksCloudflare introduces an adaptive security framework connecting risk discovery, agent governance, runtime protection, and AI-powered response in a continuous learning loop.CloudflareBuilding a post-quantum certificate authority with Merkle Tree CertificatesAs post-quantum signatures threaten to inflate TLS handshakes and certificate transparency logs, Merkle Tree Certificates offer a path to compact, auditable authentication. Cloudflare’s new certificate authority will sup…CloudflareWe tested our own WAF with frontier AI models. Here’s what we foundWe built a WAF tester that adapted each request based on what the WAF blocked or passed. This helped us explore variations that a fixed test might miss. We ran it across six attack categories on an authorized staging env…CloudflareIntroducing Threat Signals: agentic skills for open-source threat intelligence, free for every Cloudflare accountWe are expanding access to Cloudforce One's Threat Events Platform to every Cloudflare account and introducing Threat Signals. Threat Signals automatically parses open-source threat reporting, extracts structured indicat…CloudflareIs your domain using post-quantum encryption? Now you can see for yourselfCloudflare has added visibility into post-quantum (PQ) encryption in TLS 1.3 directly into HTTP Analytics, Log Explorer, and Logpush. Learn how to make sure your domain is protected with PQ encryption.CloudflareEnforce positive security with Cloudflare Application ProfilesCloudflare learns the structure of your HTTP requests and identifies deviations. You can add a positive security layer that helps reduce attack surface as AI makes it easier for attackers to generate and vary payloads.CloudflarePreventing quantum downgrade attacks against IPsecA sophisticated attacker with a quantum computer can exploit a protocol design flaw to downgrade post-quantum IPsec tunnels to classical crypto. We helped the IETF develop a transcript authentication extension to prevent…AppleThe Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language ModelsWhen language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? W…AirbnbPersonalization without user identityHow Airbnb uses proximity signals to personalize without relying on individual user history. By : Wei Jiang , Bin Xu , Bharathi Thangamani , Weiwei Guo , Sundar Srinivasavaradhan , Tracy Yu , Huiji Gao , Swapnil Ghike ,…StripeTravel’s AI dilemma at Skift Global ForumAt the “great recalibration”-themed travel conference, Airbnb CEO Brian Chesky called AI “an existential risk” to his company and “literally the best thing to ever happen” to it in the same discussion. That sums up many…SalesforceEngineering Multi-Agent AI Teams That Build and Test ThemselvesBy Sohini Arya and Manish Kumar Jha. In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Sohini Arya, AI Delivery Lead for Marketing Clou…OpenAIBasis completes a tax workbook 2x faster with GPT-6 AstraGPT-6 Astra completed a 50-tab tax workbook twice as fast as GPT-5.6 Sol, and its stronger understanding of user intent gives Basis more confidence in real-world use.OpenAIAre you a Codex Original?We’re collecting real stories of builders, tinkerers, researchers, and creators who are using Codex to do incredible things. If you want to be a part of the next chapter of the Codex Originals program, tell us more about…OpenAIThe Lenfest Institute grows landmark program with expanded OpenAI supportOpenAI is expanding the Lenfest AI Collaborative and Fellowship Program with $5 million in funding and up to $5 million in software credits and engineering support.OpenAITowards safety cases for frontier AI trainingOur early guidelines for safety cases in frontier AI training cover technical safeguards, operational practices, and investigating misalignment incidentsOpenAIHow we will do better for AustraliaOpenAI apologises for incidents involving Australian government websites and outlines stronger safeguards and support to strengthen Australia’s cyber defences.NVIDIAHow NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and EfficiencyEvery unused watt is capacity left on the table. AI factories are typically provisioned for the unlikely moment when every GPU reaches peak power, creating a...NVIDIAAdd Runtime Controls to AI Agents with NVIDIA OpenShellAI agents can be given a goal, write code, use tools, and keep working as new information becomes available. This opens the door to applications that...NVIDIANVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent MonitoringTo understand where agentic AI stands today, consider the last seismic shift in technology: the rise of the internet in the 90s. It was new and full of...Microsoft ResearchOne year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impactSince launching a year ago, the Microsoft Research Asia — Singapore lab has established a strong foundation, deepened collaboration across government, academia, and industry, and explored how frontier AI research can cre…Mistral AIMistral Opens Munich Hub to Advance Industrial AI in GermanyMicrosoftGet started with the GitHub Copilot app: a free, hands-on courseLearn to direct AI coding agents, review their work, and build reliable development workflows with this free GitHub Copilot app course. The post Get started with the GitHub Copilot app: a free, hands-on course appeared f…Hugging FaceWelcome RL Environments to the hubHugging FaceHolo4: powering generalist computer-use agentsGitLabWhat's new in Git 2.56.0?The Git project recently released Git 2.56.0 . Let s look at some of the highlights of the release, including contributions from the Git team at GitLab. What s covered: Git Merge 2026 and schedule for Git 3.0 git-history…GitHubHighlights from Git 2.56The open source Git project just released Git 2.56. Here is GitHub's look at some of the most interesting features and changes introduced since last time. The post Highlights from Git 2.56 appeared first on The GitHub Bl…GitHubHow we found 24 Android vulnerabilities using our open source AI security agentA look at the targeted AI taskflows behind these findings, the critical Android bugs they uncovered, and how to run the same open-source agent on your own app. The post How we found 24 Android vulnerabilities using our o…CloudflareIntroducing Forge: the open source pipeline for generating SDKs, CLIs, docs, and moreForge is a pluggable, open-source pipeline that runs in CI to generate SDKs, CLIs, and documentation directly from API definitions. By shifting generation upstream into individual team repositories, Forge keeps developer…CloudflareFour months of VoidZero at Cloudflare: making the open-source JavaScript toolchain faster for all humans and agentsSince joining Cloudflare, VoidZero has delivered more than 80 releases that drastically speed up JavaScript compilation, linting, and testing. From a 10x faster React compiler to Vite+ 1.0, here’s how we are building fas…CloudflareHow fast is the web? Explore billions of real-user measurements with BEACONCloudflare is open-sourcing the BEACON dataset, making billions of anonymized Real User Monitoring (RUM) performance records publicly available on Google BigQuery. Explore real-world Core Web Vitals, soft navigation metr…CloudflareIntroducing cf: the agentic CLI for the entire Cloudflare APIWe are releasing cf, our new command-line tool that mirrors the entire Cloudflare API and supports programmatic TypeScript configuration. We are also open-sourcing Forge, our internal SDK generator.AWSAWS Weekly Roundup: GPT-6 Sol and Luna, Claude Opus 5.5 on Amazon Bedrock, Strands harness, and more (September 28, 2026)If there’s one theme that defined last week, it’s choice. The frontier models keep arriving, and the interesting question is no longer just “how smart is it?” but “which model fits this step, at this cost, at this latenc…AppleFaster Rates for Federated Variational InequalitiesIn this paper, we study federated optimization for solving stochastic variational inequalities (VIs), a problem that has attracted growing attention in recent years. Despite substantial progress, a significant gap remain…PinterestPartition Finalization in Pinterest’s Next-Generation DB Ingestion FrameworkQianrui Zhang | Sr Software Engineer, Logging Platform Kanchi Masalia | Software Engineer II, Stream Processing Platform Liang Mou | Sr Staff Software Engineer, Logging Platform Yi Pan | Principal Engineer, Agent Platfor…OpenAIWayfair boosts catalog accuracy and support speed with OpenAIWayfair uses OpenAI models to improve ecommerce support and product catalog accuracy, automating ticket triage and enhancing millions of product attributes at scale.OpenAILowe’s answers millions of home improvement questions with OpenAILowe’s uses OpenAI models to power Mylow and Mylow Companion, helping customers and associates answer 25M+ questions and move home improvement projects forward.OpenAIProaction boosts sales 60% and saves 75+ hours with CodexWith Codex, GPT-Live-1, and GPT-6 Astra, Proaction builds, operates, and sells modern fleet management faster.NetflixTrading a Cloud Identity for Your Own: Workload Attestation on Managed ComputeBy Dhruv Pratap Introduction Organizations that have been around for a while usually run two identity systems side by side. One belongs to the cloud provider: IAM roles, instance profiles, execution roles. The other is y…GrabAutomated merchant review summary system with integrated feedbackIntroduction Merchant reviews contain useful details about food quality, portion size, packaging, and value. Finding these details often requires reading many comments. Aggregate ratings simplify comparisons, but they do…GitHubImproving site performance by shipping more CSSHow we fully migrated github.com away from CSS-in-JS. The post Improving site performance by shipping more CSS appeared first on The GitHub Blog .GitHubGitHub Copilot app for Beginners: How to build custom workflows with canvasesDescribe the interface you need in plain English, then let the agent build a live surface you can both use and update—so you spend less time adapting to tools and more time getting work done. The post GitHub Copilot app…DiscordDiscord Update: September 25, 2026Here s the Discord Changelog from September 25, 2026, so you can stay informed on what’s new in recent app updates!CohereCompass is coming to the cloudAnthropic2028: Two scenarios for global AI leadershipSalesforceHow Intelligent Load Shedding Prevents Cascading Failures in Tier-0 SystemsBy Ramya Subramani and Prateek Vats In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Ramya Subramani, Vice President of Engineering. R…NVIDIAEfficient MoE Training for Biological Foundation ModelsAs language models grow, scaling dense architectures becomes increasingly expensive. In a dense transformer, every token passes through every layer, so adding...MetaBringing Private Processing to Meta AI GlassesWe believe glasses are the best form factor for having AI help throughout your day. They can understand your personal context better than other kinds of devices and keep you present without picking up a mobile phone. Mos…Hugging FaceAccelerating vision-language models with LFM2.5-VL-DSparkGitHubAI-powered fuzzing with the GitHub Security Lab Taskflow AgentIn this blog post, I explain how to use the new fuzzing taskflow based on the GitHub Security Lab Taskflow Agent AI framework. The post AI-powered fuzzing with the GitHub Security Lab Taskflow Agent appeared first on The…GitHubWhen chat is the wrong UIWhat is a developer to do when they need something more tangible than a chat box? Enter canvases. The post When chat is the wrong UI appeared first on The GitHub Blog .Google DeepMindIntroducing Gemini 3.8 Live with Live AvatarAWSIntroducing enhanced custom event buses in Amazon EventBridge for enterprise-scale event-driven applicationsAmazon EventBridge announces an enhanced custom event bus for organizations scaling event-driven applications across teams and accounts. Deploy a single event bus shared across all AWS accounts in your organization, with…AppleCompressing Streaming Neural Audio Encoders via Latent-Space DistillationSystem-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the l…AppleA Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update StabilizationSemi-supervised federated learning (SSFL) trains models on clients’ unlabeled data using a teacher to generate pseudo-labels, with a small labeled seed dataset on the server. Automatic Speech Recognition (ASR) is particu…OpenAIGrab and OpenAI bring practical AI skills to Southeast AsiaOpenAI and Grab launch GO Forward with AI, a regional programme helping 30,000 partners build practical AI skills across Southeast Asia.OpenAIAirbnb widens access to GPT-6 Astra and OpenAI frontier modelsLearn how Airbnb is expanding access to GPT-6 Astra and OpenAI frontier models to help engineering teams solve bugs, design systems, and ship faster.OpenAIChatGPT Ads expands to Southeast Asia and TaiwanChatGPT Ads is expanding to Southeast Asia and Taiwan, giving eligible businesses new ways to reach people across more than 60 countries.OpenAIIntroducing MentalHealthBenchMentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.OpenAISam Altman’s remarks at the United Nations Security CouncilOpenAI CEO Sam Altman discusses AI safety, human control, and international cooperation in remarks to the United Nations Security Council.OpenAIHarvey turns legal context into stronger drafts with GPT-6 AstraGPT-6 Astra produces more structured, context-aware legal documents, freeing lawyers to focus on strategy.OpenAIHow invideo improves color grading 3x with GPT‑6 AstraWith GPT‑6 Astra, invideo plans edits with greater precision, improves color correction and grading threefold, and produces 50 custom effects in one day.OpenAIRingg’s AI agents resolve up to 65% of customer calls with OpenAIUsing GPT-5.6, Ringg powers multilingual agents across voice, chat, WhatsApp, and web for 90% less cost vs. GPT-4.1.OpenAIOpenAI extends cyber access to Ukraine for civilian defenseOpenAI is extending access to its Daybreak program to the Government of Ukraine to support the cyber defense of civilian infrastructure.OpenAITwo years of OpenAI AcademyMarking two years of OpenAI Academy and bringing AI skills to even more communities.NVIDIAHow SWE-Serve Exposes the Gap Between Local Tests and Live ServingAn AI coding agent’s patch can pass tests yet fail when the server loads a real model and handles requests. Evaluating changes to inference-serving software...NVIDIAManage Kubernetes Node Fleets with NodeWrightKubernetes manages what runs on your nodes. Managing the nodes themselves is the challenge: kernel settings, system packages, storage layouts, security agents,...NVIDIAValidate GPU Cluster Readiness Before AI Workloads LandA GPU cluster can pass every health check and still fail to run an AI workload. Even when every GPU, network link, and pod reports healthy, a 512-GPU training...NVIDIAIntroducing NV-Reason-CT Open 3D CT VLM for Radiologist Chain-of-Thought ReasoningRadiology AI has made remarkable strides in detecting abnormalities across chest X-rays, pathology slides, and 2D scans. Yet one of the most clinically rich and...Microsoft ResearchOffloaded inference for real-world physical AI roboticsRobots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanc…Hugging FaceHow to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning WorkflowsGitLabGitLab Critical Patch Release: 19.4.1, 19.3.3, 19.2.7GitHubDevelopers want more efficient software. Here’s what over 1,000 GitHub users told us they need.New research from GitHub and Yale Program on Climate Change Communication finds strong demand for tools, measurement, and practical guidance that can help developers reduce wasted compute. The post Developers want more e…GitHubRendering huge pull requests in the GitHub Copilot appHow we rebuilt the diff surface in the GitHub Copilot app to open a million-line pull request with hundreds of inline review comments. The post Rendering huge pull requests in the GitHub Copilot app appeared first on The…DropboxDropbox CTO Ali Dasdan on moving from AI adoption to transformationMeaningful AI gains depend on rethinking workflows, measuring real outcomes, and keeping human judgment at the center of how work gets done.Google DeepMindGemini 3.8 text-to-speech says helloGoogle DeepMindAdvancing Private AI Compute with secure, server-side memoryIntroducing private, server-side memory to Private AI Compute for personal AI.CohereSara HookerAppleHow to Guide Your Language FlowWe introduce a new method to guide flow matching models. Our approach, which we call probe guidance, uses the frozen internal states of an existing diffusion model to construct a guidance signal. This works using a simil…AnthropicClaude discovers a novel enzyme system with CRISPR-like repeatsStripeNew trends in global card fraud: How 3D Secure and regional mandates are affecting riskWe analyzed billions of transactions on Stripe from January 2022 to March 2026 to understand how card fraud patterns differ by region and country, what's driving those differences, and how businesses can respond.SalesforceAI Agent Observability: Making Silent Production Failures Visible and ActionableBy Archana Kumari, Yushu Yao, and Priya Lal. In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Archana Kumari, Director of Software Eng…OpenAIPriorities and principles for effective third party assessmentsOpenAI outlines priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards.OpenAIParallel cut research time and cost in half with GPT‑6 AstraGPT‑6 Astra allowed Parallel’s agents to research and synthesize labor-market data in half the time and at half the cost vs. prior models.OpenAIIntroducing GPT-6 Sol and LunaMeet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.OpenAIBetter prompt caching for GPT-6Learn how GPT-6 improves prompt caching with higher cache hit rates, new diagnostics, explicit breakpoints, and controls that reduce latency and costs.NVIDIAAccelerating a ROS 2 Node with an AI Agent and NVIDIA Isaac ROSGPU acceleration can speed up compute-intensive robotics workloads, but a fast CUDA kernel alone does not guarantee a fast ROS 2 graph. As messages move between...NVIDIAWhat’s New for Game Developers: DLSS 5 with 3D-Guided Neural Rendering, NVIDIA ACE Updates, and New RTX Kit CapabilitiesNVIDIA DLSS 5 introduces DLSS 3D-Guided Neural Rendering and granular controls that help game developers add lifelike lighting and material detail while...NVIDIATopology-Aware Workload Scheduling with NVIDIA TopographAI factories are power-limited systems that deliver maximum value when fully optimized. GPU workload placement is a key optimization. Poor workload placement...NVIDIAEnabling Private High-Performance Production AI Inference with NVIDIA Confidential ComputingAs large language model (LLM) inference increasingly processes sensitive information and proprietary model context across personal, enterprise, and regulated...Hugging FaceJun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX communityHugging FaceTransformers now runs llama.cpp quantsHugging FaceHow UK AISI and EvalEval Are Making Benchmark Results ReproducibleGitLabHow GitLab reduced code-per-agentic-flow ratio by 45%GitLab Duo Agent Platform orchestrates and automates complex tasks through agentic flows. A key part of the platform is the Flow Registry, a declarative configuration framework, built from reusable components, that compi…GitLabHow to design GitLab for enterprise scaleAt enterprise scale, even small architecture choices can have outsized consequences. A deployment that works for a handful of teams can become a constraint once thousands of developers, repositories, and pipelines depend…DiscordSafer for Teens, Same Discord for AdultsWe’re launching a revamped, privacy-focused age assurance solution including methods that won’t require an ID or selfie. Here’s how it works.CohereAI change management: A human-centric approachAWSIntroducing Amazon CloudWatch Omni: AI-powered observability for generative AI and agentic workloadsLearn how Amazon CloudWatch Omni delivers AI-powered observability purpose-built for generative AI and agentic workloads. Trace, evaluate, and experiment with AI agents across any framework—directly from your IDE or a st…AWSNow on Amazon CloudWatch Omni: collaborative AI-powered observability for your applicationsAmazon CloudWatch Omni is the next evolution of CloudWatch — unified observability that brings your applications and AI agents into one reimagined experience, with auto-discovered topology, natural language queries, and…OpenAIBuilding standards for the next phase of AIOpenAI outlines a path to shared global AI standards, calling for coordinated evaluation, reporting, and governance to improve safety.OpenAIAdvisory Group on Mathematics and Artificial IntelligenceOpenAI is working with an independent Advisory Group on Mathematics and Artificial Intelligence to guide the review and communication of emerging AI results.OpenAIHiggsfield AI ships new video features in a day with GPT-6 AstraWith GPT-6 Astra, Higgsfield AI makes video ad creation easier for small businesses and brings new creative tools to market faster.NVIDIATurn Your Latest Observations Into Timely Weather Decisions With NVIDIA Earth-2Weather-sensitive industries increasingly have access to observations that offer an earlier, more local view of changing conditions. Energy companies collect...NVIDIAHow to Evaluate AI Agents From Tool Calls to Task CompletionWhen you ship an AI agent, the key question is whether it can execute a chain of work across dozens of sequential tool calls against a live environment, and...NVIDIASimplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-TritonThe compute and memory demands of generative AI increasingly exceed what a single GPU can provide. NVIDIA TensorRT multi-device inference is a new capability...Microsoft ResearchImproving synthesis prediction of small molecules at scale with RetroChimeraCustom-made molecules are advancing medicine, materials, and agriculture, but producing them is slow and expensive. A new Nature paper highlights RetroChimera, a predictive model that helps accelerate chemical synthesis,…MicrosoftKnowledge cutoff is a poor proxy for model capabilityA model can fail on features released before its knowledge cutoff, then succeed on ones released after it. We tested hundreds of product changes and found that the date tells you far less than the work does. The post Kno…MetaInside Petal: Building the World’s First Petabit-Class Transoceanic Subsea CablePetal, the next step in Meta’s subsea innovation, will be the first subsea cable to deliver petabit capacity at transoceanic distances, connecting France and the United States over approximately 7,000 km (4,300 mi). Expe…MetaOpen-Sourcing Rebalancer: A Generic, High-Performance Library for Solving Assignment ProblemsWe re open-sourcing Rebalancer, the assignment-problem solver that has been used to solve resource allocation problems throughout Meta for over nine years. Rebalancer separates several related concerns: how to specify an…Hugging Facetokenizers v1: encode, decode and scaling, measuredHugging FacePruning LLMs Like a Physicist: Block Removal as an Ising Optimization ProblemDropboxIntroducing our new Dropbox API documentationLearn about the newly launched version of the Dropbox API documentation and its modern design and features.AWSAWS Weekly Roundup: AWS Builder Center mobile apps, Amazon Connect Talent GA, Amazon Corretto 27, and more (September 21, 2026)Living in the Netherlands, I spend a fair amount of time on trains, and that is usually where I catch up on what the builder community is writing. Until now, that meant opening a laptop or squinting at a browser tab on m…StripeAnalyzing rising fraud attempts among travel and leisure businesses on StripeLast year, Stripe data shows fraud attempts against travel and leisure businesses hit a four-year high. We analyzed payment activity from more than 200,000 active travel and leisure businesses on Stripe to understand whe…NVIDIABenchmarking LLM Inference at Scale with AIPerfYou’re deploying a model on a system. It starts up, prompts are getting responses. Now the hard question: Is this fast? Your instincts might lead you to send...NetflixLeave the Class Path in the Rearview MirrorIntroducing composable, module system native and agent friendly command line tools for modern Java development By Danny Thomas, JVM Ecosystem Team Recent work on the Java language to pave the on-ramp has made it easier t…GitLabGitLab 19.4 releasedGitLabSecuring the software factory at machine speedI joined GitLab at a moment when the way teams build and secure software has been changing rapidly. GitLab CEO Bill Staples recently framed that shift in When Code Is Abundant . When code is no longer the bottleneck, tru…GitHubShould you read the code, is RAG dead, and did Skills kill MCP?We dive into these questions and other AI hot takes on the latest episode of the GitHub Podcast. The post Should you read the code, is RAG dead, and did Skills kill MCP? appeared first on The GitHub Blog .AppleDynamically Scaled Activation SteeringActivation steering has emerged as a powerful method for guiding the behavior of generative models towards desired outcomes such as toxicity mitigation. However, most existing methods apply interventions uniformly across…AnthropicPartnering with Accenture on embedded evaluationStripeSaaS platforms are surging despite the SaaSpocalypseThe SaaSpocalypse was a useful warning for the software industry, but SaaS platforms that help businesses run core operations are more deeply embedded. New platform businesses on Stripe are up 182% year over year.PinterestBeyond Two Towers: Launching the 3-Tower Engagement Co-Train Model (Part 2)Authors: Longyu Zhao (Staff Machine Learning Engineer), Gwendolyn Zhao (Staff Machine Learning Engineer), Peng Yan (Senior Machine Learning Engineer), Yuanlu Bai (Senior Machine Learning Engineer), Yuan Wang (Senior Mach…GitLabGitLab Duo CLI takes a task from goal to doneComplex tasks hit a wall at the end of every chat turn. Developers already know what they re trying to accomplish. What stalls the work is the handoff back and forth between each step. Without a pre-defined definition of…GitLabNew MCP tools help platform teams scale automationAgentic tools are moving fast past code completion into running pipelines, opening merge requests, and triaging work on a team s behalf. The Model Context Protocol ( MCP ) has become the common way these agents reach int…GitLabSee who spent your AI credits and set fair caps per teamScaling AI across your organization depends on knowing where the budget is going and who’s using it. While a subscription cap keeps your total spend within budget, it can’t tell you how much AI was used by individual tea…GitLabOptimize your team's price-performance with hosted open weight modelsThere’s no single best model for every software development task. Implementing a new feature, diagnosing a failed pipeline, and resolving security vulnerabilities all place different demands on the model handling them. G…GitLabRate limits on GitLab.com are changingGitLab.com hosts millions of projects for teams of every size that need a platform they can rely on. Demand is climbing quickly, and we expect platform load to grow several times over this year. Predictable limits are wh…GitHubMigrating the GitHub Copilot runtime to Rust, using CopilotA rewrite this size wasn't affordable before agents. Here's what porting the Copilot agent runtime to 800,000 lines of production Rust actually took. The post Migrating the GitHub Copilot runtime to Rust, using Copilot a…AWSAWS Elastic Beanstalk introduces Cluster ModeRun an application on AWS Elastic Beanstalk Cluster Mode without provisioning or operating the compute underneath it. You provide a container image or source code; Elastic Beanstalk with service-operated compute creates…AWSNew low-cost burstable Amazon EC2 T8i instances are generally availableAWS introduces new low-cost burstable Amazon EC2 T8i instances powered by custom sixth generation Intel Xeon Scalable Processors (Granite Rapids), available only on AWS. T8i instances are among the lowest-cost EC2 instan…AppleREVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL CliffA central goal of autonomous reinforcement learning is continuous policy training without external resets. However, existing paradigms largely depend on underlying environmental reversibility, a property absent in real w…AnthropicIntroducing the Life Sciences Verification ProgramAirbnbThe guest journey, updated in real time: extending Airbnb’s sequence recommender with ChrononHow two new Chronon capabilities, Push Mode and NRT Model Transform, allows us to provide more relevant search results instantly as a guest explores, rather than waiting for the next batch run. By : Pengyu Hou , Yuli Han…YelpML based ranking using NrtsearchWe’ve extended Nrtsearch with the Inference Plugin, which embeds ML-based ranking directly in the search layer — eliminating the need for a standalone scoring service. We use Nrtsearch (read more information on the blog…SpotifyAI Changed How Spotify Builds. What We Learned (and Fixed) About Quality at Higher VelocityQuality and reliability have always been a point of pride for Spotify. We run an extraordinarily complex... The post AI Changed How Spotify Builds. What We Learned (and Fixed) About Quality at Higher Velocity appeared fi…SalesforceHow Data 360 Builds Trusted Context: The Enduring Layer for Enterprise AIBy Raveendrnathan Loganathan and Tobias Muehlbauer. How do you compile the smallest sufficient body of current, relevant, and authorized evidence for every agent turn but without loading the entire enterprise data estate…NVIDIATranslating CUDA Tile Operations from Python to Rust Using Agentic AIcuTile Rust (cutile-rs) is a tile-based system for safe, idiomatic GPU kernel authoring in the Rust programming language. Extending the Rust ownership model to...NVIDIATensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX ThorAI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answers a single prompt, an agent works through...NVIDIAHow to Use AI Agents to Prepare 3D Scenes for SimulationAgentic AI workflows can be used to prepare and validate digital twins for physical AI systems. Agents can inspect 3D scenes, author simulation-relevant data in...Mistral AIMistral x Mozilla: Private, Multilingual AI BrowsingMicrosoftYour AI coding agent evaluation is only as good as its sandboxYour AI coding agent passed the eval. But did the model know the answer, or did it find it somewhere on your machine? A correct answer can still invalidate your measurement. The post Your AI coding agent evaluation is on…GitLabWhen to use SAST versus an LLM security scannerYou re probably running some version of this experiment already: Point a frontier model at a merge request and ask it to double as a vulnerability scanner. On a single merge request, it often works well. The model reads…DiscordAdding the Discord Updates YOU Want to SeeFor Discord’s birthday, we wanted to celebrate YOU, so we asked you what you wanted to see… and we cooked up new things that you all asked for!CohereCohere & Aleph Alpha: Transatlantic Sovereign AICohereCohere and OpenText partnership announcementAWSAWS reimagines the getting started experienceAWS has reimagined the getting started experience with smart and sensible defaults to help developers get started fast so that they can focus on building. New customers can sign up and get started right away with $100 in…AppleDACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language ModelsDiffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood est…AppleGlyph: A Multi-Strategy Agentic System for Column Description and Sensitivity-Ontology Tagging of Enterprise Data CatalogsEnterprise data lakes accumulate tables faster than human stewards can document or classify them, leaving columns with missing descriptions and unassigned governance labels. This documentation debt undermines data discov…AppleHow Value Induction Reshapes LLM BehaviourConversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as helpfulness, harmlessness, and honesty. T…AppleTrajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated DistillationDiscrete flow matching generates text by iteratively transforming noise tokens into coherent language, but may require hundreds of forward passes. Distillation uses the multi-step trajectory to train a student to reprodu…AppleShared Selective Persistent Memory for Agentic LLM SystemsAgentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, discarding the configuration choices, domain constraints, data schemas, and tool-use p…StripeWhat Stripe data shows about fraud at AI startupsWe analyzed attempted fraud rates and customer abuse patterns on Stripe over the past year and found that AI companies faced 4.3x more fraud attempts than startups overall in Q3 2025.NVIDIAScaling Federated Learning Across Docker, Kubernetes, and Slurm with NVIDIA FLAREFederated learning (FL) projects often begin with a straightforward setup: one server, a few clients, and one dataset at each site. As those projects grow, the...NVIDIAHow NVIDIA NVLink 6 Delivers Multi-Layer Resiliency for AI FactoriesFor operators of large-scale AI factories, maximizing continuous output is essential for productivity. In massive-scale AI training, every GPU in the cluster...NVIDIAHow NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera RubinPower is a defining constraint for AI factories. As AI workloads demand a full compute platform to serve them, each component of that platform must maximize...NVIDIADense vs. MoE Models: Active Parameters, Throughput, and When to Choose EachHow can a 30B-parameter model activate only 3B parameters per token, and still use the capacity of the larger model? Nemotron 3.5 Lightning illustrates the...Mistral AIMaking sovereign, open-weight AI the technology frontierMicrosoftBuild an interview coach app with the GitHub Copilot SDKAn interview coach has to do more than ask questions. It needs to read a resume, follow up on an incomplete answer, and save enough context to give useful feedback at the end. Some of that work is conversation. Some of i…Hugging FaceYour Agent Aced the Task. Will It Do It Again?Google DeepMindIntroducing Gemini 3.8 Live and 3.8 Live Extended ThinkingAirbnbBeyond the model: Engineering AI infra with scientific judgementHow Airbnb’s agent harness transforms unstructured data exploration by encoding scientific methodology into scalable, reproducible, and audit-ready infrastructure. By : Wren Dougherty Ask a coding agent to analyze 100,00…SalesforceHow AI Agents Get Trusted Customer Context with Data 360 Data GraphsIn our Engineering Energizers Q A series, we explore the engineering minds behind some of Salesforce s most complex technical challenges. Today, we spotlight Alex Smith, Senior Director of AI Engineering, whose team uses…NVIDIAAccelerating Dropless MoE Training in JAX with NVIDIA Transformer EngineMixture of experts (MoE) has become one of the defining architectural trends in large-scale AI model training. DeepSeek, Qwen, and Mixtral are examples of MoE...GitLabGitLab Dedicated: Compliance for a new regulatory eraEnforcements such as NIS2 are no longer a future planning consideration. The European Union Agency for Cybersecurity s (ENISA) NIS360 report confirms that supervisory authorities are actively assessing cybersecurity matu…AWSAWS Weekly Roundup: OpenAI GPT-6 Astra on Amazon Bedrock, Amazon Quick desktop GA, Kiro for students, and more (September 14, 2026)There’s a particular energy to mid-September in New York. Pumpkin spice lattes are flowing, temperatures are dropping, and it’s nearly sweater weather. The city is back at full speed, and so is the AWS launch calendar. T…CohereWho Gets to Define the Rules for AI?SalesforceEngineering Real-Time Mobile Personalization Across iOS, Android, React Native, and FlutterIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Sneha Singla, a director of software engineering, whose team confronted a brutal engi…PinterestEvolving Pinterest’s Embedding Retrieval PlatformAuthors: Bowen Zhou | Staff Software Engineer; Shan Gao | Senior Software Engineer; Jingwen Hu | Software Engineer II; Wenjiang Chu | Staff Software Engineer The Billion-Embedding Challenge At Pinterest, the “signal” is…GitLabHow to calculate DevOps platform total cost of ownershipThere’s nothing like budget pressure to put your DevOps platform under a microscope. But subscription fees and license costs only tell one part of the story. The total cost of ownership (TCO) for a DevOps platform also i…GitLabGitLab Critical Patch Release: 19.3.2, 19.2.6, 19.1.8GitHubMarketing ops as code: Automating events from planning to follow-up on GitHubIf you can write down how you do your work, you can automate it. Here's what I did to support GitHub's APAC marketing team. The post Marketing ops as code: Automating events from planning to follow-up on GitHub appeared…ApplePutting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question AnsweringEvaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from th…AppleDiscoSign: Discourse-Aware Text to Sign Language Gloss TranslationSign language processing systems have traditionally operated at the sentence level, ignoring critical discourse phenomena fundamental to sign language comprehension. We introduce DiscoSign, a computational approach for d…AppleSimpleDesign: A Joint Model for Protein Sequence and Structure CodesignProteins are fundamental to biological processes, with their function determined by the complex interplay between the amino acid sequence and the three-dimensional structure. Developing generative models capable of under…PinterestBuilding Pinterest’s VLM Serving Stack on NVIDIA DynamoLei Pan | Senior Software Engineer; Salina Wu | Senior Software Engineer; Cristian Lopez | Software Engineer I; Guangtong Bai | Staff Software Engineer; Soam Acharya | Principal Engineer; Saurabh Vishwas Joshi | Principa…NVIDIAFrom Wafer-Out to First Token: Codifying Supply Chain Expertise with Nemotron and Palantir FoundryNVIDIA has one of the largest and most complex supply chains in the world, and its performance is measured from wafer-out to first token. The interval is in two...NVIDIAHigh-Throughput Structure Prediction with BioNeMo Inference RuntimeBiomolecular structure prediction is now often run at proteome scale, where the goal is to move an entire worklist through the pipeline efficiently. NVIDIA...NVIDIAHow Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 UltraDeploying a large language model is only the first step toward production-ready serving. Production teams also need to serve as many concurrent users as...Mistral AICloudera and Mistral Partner to Bring Specialized, Sovereign Intelligence to Enterprise DataLyftRefreshing the Travel-Time Map Behind Lyft’s Marketplace: Rebuilding Neighborhood Reachability…Refreshing the Travel-Time Map Behind Lyft’s Marketplace: Rebuilding Neighborhood Reachability Signals Every time Lyft calculates pricing to balance a market, nudges a driver toward an under-served pocket of a city, or p…Hugging FaceRebuilding AUTOMATIC1111 with Gradio WorkflowHugging FaceAsync GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCLGitLabPrepare for the Cyber Resilience Act's 24-hour reporting deadlineStarting on September 11, 2026, many businesses that place software on the European Union (EU) market will have 24 hours to file a report once they learn that a vulnerability in one of their products is being actively ex…GitLabCo-Create: Building GitLab with our usersGitLab users bring firsthand experience of the workflows they want to improve. Through our Co-Create program, they collaborate directly with us to design, build, and deliver product improvements that benefit more teams.…GitHubGitHub Copilot app for Beginners: Using the diff, terminal, and browserChecking agent-generated code usually means hopping between tabs. Learn how to view diffs, run terminal commands, and preview web apps side by side in the GitHub Copilot app. The post GitHub Copilot app for Beginners: Us…GitHubGitHub availability report: August 2026In August, we experienced five incidents that resulted in degraded performance across GitHub services. The post GitHub availability report: August 2026 appeared first on The GitHub Blog .CohereIntroducing North Small Translate: A leading sovereign open-weight machine translation modelSalesforceEnterprise AI Accuracy: Building a More Trustworthy RAG ApplicationBy Sivakumar Shanmugam, Kartik Muktinutalapati, and Palani Ramanathan. Your retrieval-augmented generation (RAG) application passes its benchmarks, reaches production, and begins confidently giving customers the wrong an…NVIDIACUDA Toolkit 13.4 Adds Windows on Arm Support and Greater Control over Shared GPUsEvery NVIDIA CUDA Toolkit release adds functionality and performance improvements that help developers get more from NVIDIA GPUs and the broader NVIDIA software...NVIDIAWhen to Use Encode-Prefill-Decode Disaggregation to Accelerate Multimodal Model ServingEncode-prefill-decode (EPD) disaggregation is an inference optimization technique for multimodal models that separates the vision encoder stage from the prefill...Mistral AIModernizing complex legacy code with AI agents.MicrosoftYour work might not need the smartest modelThe smartest model can cost five times more and deliver the same result, or even a worse one. See how evaluating your own work helps you get more value from your agent budget. The post Your work might not need the smarte…Hugging FaceIBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly licenseAWSIntroducing Amazon EBS Volume Clones across AWS accountsAWS introduces Amazon EBS Volume Clones with cross-account copy, so you can create copies of your EBS volumes into other AWS accounts and optionally re-encrypt them with an AWS Key Management Service (AWS KMS) key in the…AnthropicAn alignment assessment of recent cybersecurity incidentsSpotifyWhy Spotify Is Not Using Bayesian A/B TestingClearing the confusion about what Bayesian A/B testing is. The post Why Spotify Is Not Using Bayesian A/B Testing appeared first on Spotify Engineering .SalesforceHow Metadata and Runtime Telemetry Reveal Which Enterprise System Problems to Fix FirstBy Karishma Lalwani, Anand Vardhan, and Sanjeevani Bhardwaj. In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Karishma Lalwani, Vice P…NVIDIAIntroducing CUDA Rust: Two Tracks for Writing GPU KernelsIn September 2026, NVIDIA announced it is leaning into native GPU programming in Rust. CUDA C++ and CUDA Python are mature, enterprise-grade toolchains, and...Hugging FaceSafety for Whom? Refusing the Right Subset of a Topic, Not the Whole TopicGitLabBring your own model to GitLab Duo Self-Hosted with Microsoft FoundryFor many organizations, the question about AI coding tools isn t whether they help, but where the code goes. Teams under data sovereignty, residency, or regulatory constraints need to know which network handles their sou…GitLabGPT-6 Astra on GitLab: Faster runs, fewer tokens usedOpenAI s newest frontier model GPT-6 Astra is now on GitLab Duo Agent Platform, delivering faster runs and lower token usage. In GitLab s internal evaluation, GPT-6 Astra finished a typical run 43.4% faster than GPT-5.6…DiscordDiscord Patch Notes: September 8, 2026Check out the finer details of the more technical fixes implemented into Discord recently.Google DeepMindAlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genomeAlphaGenome Atlas maps the molecular effects of 9 billion single-letter DNA variants across the human genome.CohereInside the megakernel serving engine for North Mini CodeAWSAWS Weekly Roundup: Claude Fable 5.1 on AWS, Amazon Linux 2027 preview, AWS Certified AI Business Strategist, and more (September 7, 2026)Last week, Claude Fable 5.1 became available on AWS. According to Anthropic, Claude Fable 5.1 delivers frontier intelligence for ambitious tasks across coding, scientific research, and enterprise workflows. Claude Fable…AnthropicIntroducing Claude Design by Anthropic LabsNVIDIAFrontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA JetsonRunning reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run...NVIDIABuilding a Memory-Driven Agent with NVIDIA NemoClawEnterprise work spans messages, decisions, projects, and obligations that change over time. An AI agent that starts without this context must reconstruct it...GitHubProject HydraFusion: Frontier quality via multi-model orchestrationIn controlled offline evaluations, HydraFusion’s selective coding workflows matched or exceeded the evaluated Opus 5 baseline while reducing estimated workflow cost. Now available as a research preview in GitHub Copilot.…CohereAidan PeppinAnthropicFormalizing Fermat's Last TheoremSpotifyPortal by Spotify cut my Claude Code token usage by 90%Most of what an AI coding agent does for me isn't thinking. It's I/O. The post Portal by Spotify cut my Claude Code token usage by 90% appeared first on Spotify Engineering .SalesforceHow Agentforce Achieves 100% Deterministic Rendering for AI Agent UXIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Abhi Rathna, Product Management Director for Agentforce, whose team tackled a fundame…MetaZGateway: Learnings from Putting a Proxy in Front of ZippyDBWe’re introducing ZGateway, the proxy we are using to unify traffic through ZippyDB, Meta’s most widely-used key value store. As a bonus, it also enables admission control, load balancing, cross-region resilience, and ri…InstacartAgentic Machine Learning Modeling at InstacartTilman Drerup, Moe Moazzami, Shih-Ting Lin, Greg Reda (and many more) Introduction At Instacart, artificial intelligence is fundamentally changing the way our machine learning engineers operate. In a prior blog post , we…Hugging FaceTraining a coding model to paint watercolours with TRL and OpenEnvHugging FaceGive Your Coding Agents a Memory You OwnHugging FaceFine-tuning a 350M Model for Better Structured Outputs in 100 GRPO StepsHugging FaceNeoMME: an efficient Multimodal-native and Multilingual EncoderGitHubGitHub Copilot app for Beginners: Run several agents at onceLearn how to run parallel agents in the GitHub Copilot app, and experience the moment it stops feeling scary and starts feeling powerful. The post GitHub Copilot app for Beginners: Run several agents at once appeared fir…Google DeepMindIntroducing WeatherNext 3, our most advanced and accurate global weather AI modelCohereAutomation’s Early Footprint: The ATE DatasetCohereDana ArsovskaMetaAn Organizational Second Brain: Building an AI That Learns From ExpertsWe’ve built an AI agent that acts as a secondary expert for a given domain, making deep specialist knowledge readily available and preserved for anyone in an organization to access, share, and build upon. This is not a t…Hugging FaceReal-Time Intelligence with IBM Time Series Models on ConfluentGitLabCritical remote code execution in vm2, a widely used Node.js sandbox libraryGitLab s Threat Research Group found a critical sandbox escape vulnerability in vm2, one of the most widely adopted Node.js sandboxing libraries. The vulnerability uses a configuration copied straight from vm2 s own READ…GitLabGitLab’s internal playbook to foster AI-fluent technical teamsGive two engineering teams the same AI tool and you can end up with two very different outcomes. One team ships faster with fewer bugs, while the other gets burned by an agent that confidently generates the wrong output.…GitHubDecoding the new AI lingo: Loops, harnesses, squads, hill climbing&#8230; oh my!From loop engineering to harnesses, squads, and open weights, the GitHub Podcast breaks down the AI terms showing up in developer conversations. The post Decoding the new AI lingo: Loops, harnesses, squads, hill climbing…GitHubHow we make AI coding more cost efficient without sacrificing task qualityWhy shorter outputs can cost more, and how GitHub Copilot reduces wasted work across the complete coding task. The post How we make AI coding more cost efficient without sacrificing task quality appeared first on The Git…Google DeepMindIntroducing Gemini 3.8 Flash and 3.8 Flash CyberGoogle DeepMindProactive cyber defense for governments and enterprisesCohereThe Benefits of Small Language Models (SLMs)AppleREFACTOR-VLA: Unsupervised Library Learning of Typed Motor ProgramsMost current vision-language-action (VLA) models—such as OpenVLA, π0, RT-2, and RDT-1B—are “monolithic.” This means they generate raw motor commands or very short sequences of actions, without organizing behaviors into r…PinterestBecoming an AI TeamJohn Grass | Sr. Manager, Engineering A Fundamental Transformation An AI team is fundamentally more than just a group whose members incorporate AI tools into their existing workflows. The journey to becoming an AI team n…Hugging FaceIntroducing @huggingface/kernels: 200+ WebGPU Kernels for Local AIHugging FaceBenchMIRT: What are LLM benchmarks actually measuring?Google DeepMindIntroducing agentic video understanding with GeminiAnthropicDeveloping Enterprise Frontier Safeguards with our customersAnthropicDeveloping Enterprise Frontier Safeguards with our customersAnthropicClaude Fable 5 and Claude Mythos 5AnthropicDeveloping Enterprise Frontier Safeguards with our customersSalesforceWhy AI Agents Get the Right Facts but the Wrong Answer—and How GraphRAG HelpsBy Sivakumar Shanmugam and Indrojit Deb.An AI agent rejects a valid customer return. You check its citation: the policy really does say 90 days. You check the purchase: it falls outside that window. Both facts are correc…Microsoft ResearchGigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation modelsWhat if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. T…LyftRerouting the Stream: How Lyft Moved to the Apache Flink OperatorWritten by Maheep Myneni , Arda Kuyumcu , and Prem Santosh Udaya Shankar at Lyft. Why We Migrated : Technical Debt Meets Modern Streaming Demands Over the past several quarters, Lyft’s Streaming Compute team retired our…DropboxTesting cookie behavior across hundreds of web surfaces with our in-house auditorOur cookie auditor acts like a privacy-conscious user by visiting web pages and checking that they only load cookies consistent with that user’s preferences.AWSAWS Weekly Roundup: Welcome DuckLabs to the team, Agentic Resource Discovery (ARD), and more (August 31, 2026)The news that interested me the most last week was the DuckLabs acquisition. AWS has signed a definitive agreement to acquire DuckLabs, the Amsterdam-based company behind DuckDB, the popular open source analytical databa…AWSAmazon EC2 R9g and R9gd instances powered by AWS Graviton5 processors are now generally availableAmazon EC2 R9g and R9gd instances powered by AWS Graviton5 are now generally available, delivering up to 25% better compute performance than R8g, ideal for databases, in-memory caches, and real-time analytics.AnthropicImproving our alignment and security practicesSalesforceHow AI-Powered Attacks Led Salesforce to Reinvent Hyperscale DDoS DefenseIn our Engineering Energizers Q A series, we spotlight the engineering leaders driving innovation across Salesforce. Today, we meet Herman Kwong, EVP of Software Engineering, whose team developed DREAM (DDoS Response and…NetflixMAPS: Netflix’s Multimodal Asset Personalization at ScaleBy Emma Yanyang Kong , Aditya Deshpande , Asad Abbasi , Bowei Yan , David Fagnan , Ashish Rastogi , Dhaval Patel , Ray Zhang Introduction The Netflix experience is a journey of discovery. Every visual cue, from the artwo…Hugging FaceThe Open ASR Leaderboard Adds Its First Global South LanguageGrabData Mesh at Grab (Part III): Operationalizing data reliability with automated DPIsIntroduction In the first two parts of this series , we described how Grab approaches data mesh through the Signals Marketplace: a way for teams to publish, discover, and reuse trusted data products across domains. Part…CohereGenerative AI for business: Use cases, benefits, and adoptionAppleAgent Seer: Synthesizing Scenarios from Specification UnderstandingEvaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain e…AppleLLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Probabilistic BeliefsModern AI systems are being deployed in complex domains such as medicine, science, and law, where there is often not a single correct answer given the observed evidence. Such systems must be able to represent and update…GitLabHow to recognize your team with GitLab AchievementsEvery team runs on people who go above and beyond. The engineer who fixes the flaky test nobody else will touch. The reviewer who turns your merge request around in an hour. The teammate who finishes their certification.…GitLabGitLab compliance frameworks: Adhere to SOC 2 in minutesCompliance is the part of software delivery that everyone agrees is important, yet nobody enjoys. It often lives in spreadsheets, screenshots, and the quiet dread of an upcoming audit. GitLab s custom compliance framewor…GitHubOpenClaw went viral. Meet the maintainers building and securing it.OpenClaw is the fastest-growing project in GitHub history. Peter Steinberger and several maintainers share what they learned in the project's first six months. The post OpenClaw went viral. Meet the maintainers building…Google DeepMindPiloting the world's first double-blind AI evaluationsPiloting the world's first double-blind AI evaluationsGoogle DeepMindGemini Omni 1.1 Flash lets you build with more controlCohereParse Brings Enterprise Intelligence at ScaleCohereFDEs: Building Capacity, Not DependencyAppleFrom Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge AnswersDesigning effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holi…AnthropicExpanding our support for scientistsAnthropicPreviewing the Model Hardware StandardPinterestScaling Conditional Learned Retrieval for Pinterest Home FeedDevin Kreuzer | Sr. Machine Learning Engineer; Yichi Wang | Machine Learning Engineer I; Sujan Reddy Ale | Machine Learning Engineer I; Zelun Wang | Sr. Machine Learning Engineer; Hongtao Lin | Sr. Machine Learning Engin…Hugging FaceTraining and Finetuning Multi-Vector Embedding Models with Sentence TransformersGitLabMaking room for what's next in the GitLab UIThroughout this year, the product interface has been in a season of reduction. On the heels of dark mode , the tide has been moving out with a quieter application chrome, overall color reduction, and neutral controls. It…GitLabGitLab Patch Release: 19.3.1, 19.2.5, 19.1.7GitLabGit was built for humans — agents need an upgradeThe industry is now racing to rebuild source code management for agents. We showed our answer at GitLab Transcend , but let’s reiterate why rebuilding the Git backend is only half the problem. Three things break when age…GitHubGitHub Copilot app for Beginners: Automate Dependabot pull request triageManaging library updates can be tedious at times. Learn how the GitHub Copilot app can handle this type of repetitive task. The post GitHub Copilot app for Beginners: Automate Dependabot pull request triage appeared firs…Google DeepMindIntelligent transcription with Gemini 3.5 TranscribeNow you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.AppleLuce: Relightable Gaussians for 3D Asset GenerationHigh-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. To support relighting and integration into standard rendering pipelines, the representation should include phy…AppleIDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model PretrainingRecent advancements in large language models have intensified the need for efficient and deployable models within limited inference budgets. Structured pruning pipelines have shown promise in token efficiency compared to…ApplePROOF-Gen: From Optimized Data to Better DistillationSupervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run…MicrosoftTry Azure SRE Agent with no always-on chargesWe are happy to announce a 30-day trial experience for Azure SRE Agent. New customers can create and configure the SRE Agent at their own pace, with no charges for setup time or keeping agents ready. During the trial, yo…MicrosoftStart here: Azure SQL Foundations seriesMost developers I talk to aren t asking whether Azure SQL Database can handle their next app. They re asking where to start when it comes to modernization, migration, and AI in the database. If you re reading this, you r…Hugging FaceWire It, Run It, Deploy It: AI Workflows in GradioHugging FaceQuantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision originalHugging FaceGranite 4.2 LLMs: How They're BuiltGitLabScale software delivery pipelines in isolation without owning the runner fleetMany enterprises choose GitLab Dedicated for a clear reason: a secure and compliant, single-tenant GitLab instance, managed by GitLab. As agentic workflows drive more pipeline volume, full data isolation and runner infra…GitHubHow to evaluate LLMs before productionThese are the lessons we learned evaluating LLMs for real-world secret scanning. The post How to evaluate LLMs before production appeared first on The GitHub Blog .CohereThe state of sovereign AI adoption in 2026AWSHappy 20th Birthday, Amazon EC2On the 20th Anniversary, we recognize how AWS has continued to push the boundaries of what cloud computing can deliver, building custom silicon for general-purpose and AI workloads and expanding EC2 into new form factors…AppleSTARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal GenerationUnified multimodal models that understand, reason over, and generate interleaved text–image sequences remain structurally fragmented: existing approaches either sacrifice visual fidelity through discrete tokenization, im…AnthropicFunding better evaluations of AI’s impact on wellbeingAirbnbProject Lighthouse — Part 3: Introducing project-lighthouse-anonymizeProject Lighthouse — Part 3: Introducing project-lighthouse-anonymize The data in Project Lighthouse is powered by privacy-preserving anonymization code. We’ve put this code into open source, and published two new techni…SalesforceFrom Prediction to Action: How to Turn AI Outputs Into DecisionsBy Ali Nahvi, Akshit Behera, Yvonne Fan, and Melissa Ramey. Your models are accurate. Your pipelines are healthy. Your dashboards are working. And your system is still failing its users. Imagine thousands of dashboards a…MozillaIntent to Ship: JPEG XLIt isn t often that new image formats land in browsers. In the early 2000s we had JPEG, GIF, and PNG. The 2010s gave us WebP, which was a modest step up from JPEG. But the 2020s have given us two new image formats that a…Mistral AIMistral x HUMAINMetaMTIA 300: Meta&#8217;s First Training Chip with Built-in NICs and Communication-Offloading EnginesMTIA 300 is the first of Meta’s family of in-house training and inference accelerators optimized for training ranking and recommendation models. We’re sharing how MTIA 300’s built-in NIC chiplets allow it to meet the com…MetaMetaRoCE: A New RDMA Transport Built for AI-Scale EthernetTraining and serving frontier AI models depends on fast, reliable networks that move data between GPUs without wasting compute cycles. To meet this challenge at scale, Meta designed MetaRoCE – a clean-sheet RDMA transpor…GitLabWhen code is abundantI returned from the holiday break in January convinced that something fundamental had changed. Large language models had reached the point where they could produce useful code reliably enough, and cheaply enough, to chan…GitHubYour alt text passes automated checks. That doesn&#8217;t mean it&#8217;s any good.We built a plugin for the GitHub Accessibility Scanner to make sure your alt text is actually accessible. Here's how it works. The post Your alt text passes automated checks. That doesn t mean it s any good. appeared fir…AWSAWS Weekly Roundup: Student Rewards on AWS Builder Center, Local Zone in Las Vegas, and more (August 24, 2026)During my time at AWS, I have always looked for opportunities to work with students. I have delivered over 50 talks at universities across the region, and watching the potential in the room is always a strong motivator.…AppleBeyond Visual CoT: Internalized Visual Thinking for Proactive Video ReasoningMultimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an int…CohereTiny Aya VisionNetflixA Tale of Two Flink AutoscalersSamuel Yeboah , Francesco Di Chiara and Mingliang Liu Today, Netflix runs two Flink autoscalers. That is exactly one more than we want. We built the first one in-house years ago, when there was no mature option suited to…Hugging FaceMeasuring benchmark optimization in speech recognitionHugging FaceHow Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with CodeGrabBuilding Jarvis Pro: Route first, answer laterIntroduction The first Jarvis Pro prototype could produce answers that sounded right. That was the problem. One early answer looked polished: it named the merchant, summarized the week, and recommended pushing promotions…Google DeepMindFrom Atari to EVE Online: Building on 15 Years of AI Research in GamesGoogle DeepMind partners with game studios to prototype breakthrough AI gameplay.AWSAWS Glue 6.0 now available with 30% lower price and full Apache Iceberg v3 supportAWS Glue 6.0 is built on a fully modernized runtime, Apache Spark 4.1, Python 3.13, and Scala 2.13, delivering 30% lower pricing than previous AWS Glue versions.YelpBuilding Menu Vision: Real-Time Dish RecognitionMenus aren’t just lists, they’re a window into a restaurant’s unique offerings, specialties, and personality, shaping where and what we choose to eat. But here’s the challenge: reading “Kung Pao Chicken - Stir-fried chic…StripeFive monetization trends from global pricing leadersAs AI transforms software economics, the standard revenue playbook is breaking down. Learn how leaders around the world are preparing for agent buyers, updating processes for faster pricing iteration, and building more f…SalesforceWhy AI-Generated Code Is Easy but Engineering Trust Is HardImagine an AI coding agent generating a large pull request filled with AI-generated code. Everything looks exactly as it should: Later, your team discovers the agent quietly misinterpreted an ambiguous requirement near t…Microsoft ResearchBroadening access to Skala creates a faster path to predictive DFTSkala 1.1, the updated deep-learning exchange-correlation functional from Microsoft Research, provides greater accuracy, expanded accessibility across the computational chemistry ecosystem, and a living benchmark to trac…Mistral AIAgentic Search. More accurate and efficient results from your AI systems.Hugging FaceUp to 3.2x Faster Inference with LFM2.5-DSparkGitLabGitLab 19.3 releasedGitLabWhen your backlog outgrows your team, GitLab scales remediationSecurity teams have historically struggled to keep up with triage and remediation when development was happening at human speed. Today, that challenge is exacerbated by developers writing and shipping code at machine spe…GitLabRun agentic software delivery inside the boundaries you already trustMany enterprises choose GitLab Dedicated for a clear reason: a single-tenant instance, managed by GitLab, in a cloud region they select. That isolation already covers source code, project data, and the software developme…GitLabBuild custom flows in minutes with the Flow Creator agentCustom Flows already let teams turn manual, multi-step work into automation that runs on GitLab events. But writing one meant learning the Flow Registry schema first. That requirement is a real barrier. The people who un…GitHubThe August 17 outage, and the work aheadAn update on the August 17 outage and the steps we're taking to improve reliability. The post The August 17 outage, and the work ahead appeared first on The GitHub Blog .CohereCultural Awareness in Global AIAppleScaling Laws for Mixture Pretraining Under Data ConstraintsAs language models scale, the amount of data they require grows – yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size. A common strategy is to mix this scar…AppleProgressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASRCode-switching (CS), alternating languages within the same utterance, poses significant challenges for automatic speech recognition (ASR) due to limited CS training data. This paper applies an iterative pseudo-labeling t…AppleMultilingual Knowledge Transfer under Data Constraints via Lexical InterventionsCross-lingual knowledge transfer is critical for building high-performing multilingual language models for languages with insufficient training data. When target language data is scarce, the knowledge required for many d…StripeWhy global workers are driving demand for stablecoin payoutsPlatforms like DoorDash, Meta, and Deel already enable stablecoin payouts for global workers. We surveyed 2,300 workers in 20 countries to see what's driving stablecoin demand, where the opportunity is highest, and how o…Hugging FaceLFM2.5 Q4\_0 Checkpoints from Quantization-Aware DistillationGitLabFrom chaos to context: Building an AI dev workflowThere s a particular kind of frustration that happens when prompting an AI assistant with the same correction multiple times in a single session. The marvels of modern large language models (LLMs) make it so you re worki…GitHubGitHub Copilot app for Beginners: Managing your workIf you’re juggling multiple Copilot sessions, use the My work pane to track what's in flight, what's done, and what's next. The post GitHub Copilot app for Beginners: Managing your work appeared first on The GitHub Blog…AppleThe P-Completeness of Inverted Index Traversal: On the Complexity of Evaluating Boolean Query DAGsModern AI agents increasingly rely on search infrastructure to execute complex, neuro-symbolic reasoning workflows. These workflows often compile into deeply nested, non-monotonic Boolean queries over text fields. Howeve…AppleExamining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System PromptsLarge language models (LLMs) exhibit a wide range of human-like behaviors, from expressing thoughts and emotions, to engaging in relationship-building with users, to refusing requests and maintaining boundaries. Despite…AirbnbHow we knew COVID was over (and what our models had to unlearn)When we retrain, when we rebuild, and when we leave a model alone. By: Harrison Katz A forecast that carries weight The Forecasting Data Science team at Airbnb produces many of the forecasts the rest of the company plans…Hugging FaceMulti-Vector (Late Interaction) Embedding Models with Sentence TransformersHugging FaceHow Much Memory Does Your Agent Actually Need?GitLabFrom OpenTofu to Argo CD: GitLab as your AWS control planeConfiguring cloud-based environments is complex, as it requires considering networking, subscriptions, services, and all the components that make up the underlying infrastructure. Manual setup is error-prone, inconsisten…GitLabAvoid the massive end-to-end tax of default full history clonesIt s easy to think of git clone as a client-side operation, but the settings of this operation impact the server side and all networks in between. When you run a default full history clone, the server has to walk the ent…DropboxImproving infrastructure efficiency for growing demand in the age of AIAs demand for AI continues to grow, so does the infrastructure needed to support it.AWSIn the works: AWS Builder Lofts in Berlin, Hyderabad, and São PauloEach location will be a permanent community space to offer free workshops, networking events, pitch nights, content creation spaces, collaboration/co-working areas, and event hosting for developers, students, or tech pro…AppleGRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual SettingsReinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models bu…AppleMVICAD2: Multi-View Independent Component Analysis with Delays and DilationsMachine learning techniques in multi-view settings face significant challenges, particularly when integrating heterogeneous data, aligning feature spaces, and managing view-specific biases. These issues are prominent in…AppleA Specialized Semismooth Newton Method for Kernel-Based Optimal TransportKernel-based optimal transport (OT) estimators offer an alternative, functional estimation procedure to address OT problems from samples. Recent works suggest that these estimators are more statistically efficient than p…StripeNew currency capabilities for global businesses to cut FX costsTwo product upgrades make it easy for global businesses to manage FX entirely on Stripe. We’re expanding multicurrency settlement to more markets and currencies, and we’re introducing the ability to convert currencies in…SalesforceHow to Evaluate Production AI Agents: Measure System Outcomes, Not ConversationsImagine deploying a production AI agent that confidently tells a customer, I have processed your refund. The customer leaves satisfied. The conversation looks perfect. Your monitoring shows no obvious errors. Then someon…Hugging FaceSame Cluster, 33 Points More Utilization: What Changed Was the OrderGitLabCritical remote code execution in Serena, a popular MCP coding agentSerena, one of the most widely used AI coding agents, ran attacker-supplied code the moment a developer opened a project. GitLab s Threat Research Group found a critical server-side template injection ( GHSA-pp25-4cg4-qc…GitLabGitLab Critical Patch Release: 19.2.4, 19.1.6, 19.0.8, 18.11.11GitHubHow canvases make agentic workflows visible, steerable, and cost-efficientChat is great for intent, but agent work gets lost in the scroll. Here is how I use canvases with my agentic workflows—and why your workflow also deserves a canvas. The post How canvases make agentic workflows visible, s…DiscordA letter to the Discord community in BrazilOn the order from Brazil s ANPD to suspend screensharing and video features in the country, how we re complying, and our commitment to our community in Brazil.AWSAWS Weekly Roundup: EC2 application status checks, IAM role manager, OpenAI Daybreak on Bedrock, and more (August 17, 2026)Last week, AWS contributors joined the OpenSearch and Valkey communities at Open Source Summit Korea 2026 and MCP DevSummit Seoul 2026 to meet open source developers and contributors. At the four-day event, community lea…Hugging FaceState of Open Models: Summer 2026 ObservationsGitHubHow to bring your software delivery workflow into GitHub with agent appsSee how four GitHub agent apps can help you scope, secure, roll out, and ship a feature across the SDLC–all without leaving GitHub. The post How to bring your software delivery workflow into GitHub with agent apps appear…AnthropicHow Claude's text watermarking worksAnthropicHow Claude's text watermarking worksSpotifyWhen Can LLMs Replace Humans in A/B Tests?TL;DR: LLM predictions can stand in for human outcomes in A/B tests, but only by assumption, not by design.... The post When Can LLMs Replace Humans in A/B Tests? appeared first on Spotify Engineering .Hugging FaceWhat We Learned by Reproducing 2,200 papers from ICMLHugging FaceRecord, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage BucketsGitLabHow I built a demo generator with GitLab Duo Agent PlatformA demo used to take me days to build — screenshots, narration, stitching it together in an external tool, chasing feedback — and every time the feature changed I d have to start over. A few months ago, I handed most of t…GitHubWhat 50 open source projects taught us about security in the AI eraSee how the open source projects in Session 4 of the GitHub Secure Open Source Fund combined AI-assisted workflows, maintainer expertise, GitHub security tools, expert guidance, and funding to improve project security. T…GitHubYour guide to GitHub Universe 2026 is here: The schedule just launched!The GitHub Universe session catalog is live. Explore interactive workshops, community talks, demos, and panels. Plus, register before August 19 to save $300. The post Your guide to GitHub Universe 2026 is here: The sched…Google DeepMindIntroducing Gemini 3.7 FlashAppleWhen Unlearning Is Free: Leveraging Low Influence Points to Reduce Computational CostsAs concerns around data privacy in machine learning grow, the ability to unlearn, or remove, specific data points from trained models becomes increasingly important. While state of the art unlearning methods have emerged…SalesforceHow Agentforce-Powered AI Security Workflows Accelerate Incident ResponseIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Zac Taylor, a software engineering architect on the Trusted Services team. Zac and hi…Microsoft ResearchMindTopo reveals VLMs&#8217; spatial reasoning abilitiesA path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLM…MetaHow We’re Building Scam Alert on WhatsApp With End-to-End Encryption and Verifiability GuaranteesWhatsApp is committed to helping people stay safe while protecting the privacy of their messages. As scam tactics evolve — from impersonation to social engineering to AI-generated lures — we re always evolving as well, s…Hugging FaceLFM2.5-VL-3B for Better and Faster Vision Capabilities for the EdgeGrabGrab Bench: Evaluating AI on Grab-shaped production workIntroduction What worried us wasn’t the hallucination, it was the subtle plausibility. Answers an engineer could easily read past and accept: a right-looking Structured Query Language (SQL) query, a plausible tool call,…GitLabHow GitLab tracks vulnerabilities through refactors and reformattingEvery day, security scans face the same problem: an agent or a developer adds a comment, reformats a file, or moves a function, and a naive vulnerability tracker suddenly reports the same finding twice. Security teams en…GitLabA sandbox is only as closed as what an AI agent can reachIn July, OpenAI and Hugging Face responsibly disclosed an OpenAI model under internal evaluation escaped its sandbox, reached the open internet, and accessed Hugging Face’s internal production infrastructure. The agent t…GitLabGitLab Patch Release: 19.2.2, 19.1.4, 19.0.6GitHubYour contributors are AI-first now. Is your project?AI contributors are already in your queue. AutoGPT maintainer Nicholas Tindle shares the repo instructions, gates, and boundaries that keep maintainers in control. The post Your contributors are AI-first now. Is your pro…GitHubGitHub availability report: July 2026In July, we experienced eight incidents that resulted in degraded performance across GitHub services. The post GitHub availability report: July 2026 appeared first on The GitHub Blog .GitHubGitHub Copilot app for Beginners: Write your first promptLearn how to write your first prompt in the GitHub Copilot app, choose the right context and model, and start your first task with confidence. The post GitHub Copilot app for Beginners: Write your first prompt appeared f…Google DeepMindPutting sign language AI into users’ handsIntroducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.AirbnbFlexible Authentication: Reimagining authentication for millions of users at AirbnbRebuilding login and signup surfaced product insights, not just technical challenges. Here’s how we designed Flexible Authentication at the intersection of product intuition and technical architecture. By : Jose Santos ,…StripeMapping the AI economyAI companies are undergoing rapid global expansion while achieving unprecedented rates of growth. We analyzed Stripe data to understand where global demand is the strongest, and how companies can build to best capture th…SalesforceHow Standardizing Product Telemetry Reduced Time to Insight by 97%By Rounak Mehta, Adrian Eng, Sandeep Singh, and Kishore Kanchapalli. In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Adrian Eng, Dire…Microsoft ResearchIntroducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented MeasurementRadiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introdu…Mistral AIIn-region inference, open models, and new European infrastructure for sovereign AI.GitHubFrom coder to orchestrator: How agents shift the role of a developerDevelopers are owning more of the delivery system around code, not just code itself. Join us during GitHub Universe to meet other devs, learn something new, and explore what's next. The post From coder to orchestrator: H…DiscordDiscord Update: August 11, 2026 ChangelogHere s the Discord Changelog from August 11, 2026, so you can stay informed on what’s new in recent app updates!DiscordLink Discord and WoW to Keep the Guild Chatting Wherever You AreGather your guildmates! Link your Discord and Battle.net accounts to unlock synced guild chat across Discord and World of Warcraft.GitHubUsing the GitHub Copilot SDK for JavaEnterprise Java developers have a new superpower—drive GitHub Copilot from idiomatic Java code with annotations, virtual threads, and more. The post Using the GitHub Copilot SDK for Java appeared first on The GitHub Blog…AWSAWS Weekly Roundup: AWS Heroes Summit, Web Search on Amazon Bedrock, Dogwood, Kiro Crew, and more (August 10, 2026)Last week, we brought together AWS Heroes from around the world to connect, collaborate, and celebrate the builders who go above and beyond for the AWS community. The AWS Heroes Summit, an invite-only annual gathering, b…NetflixHow and Why Netflix Built a Real-Time Distributed Graph: Part 3 — Querying the graph with gRPC…How and Why Netflix Built a Real-Time Distributed Graph: Part 3 — Querying the graph with gRPC execution API Authors: Nilesh Mishra and Ajit Koti This is the third entry of a multi-part blog series describing how we buil…AppleBeyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language ModelsLarge Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks, including document processing and code generation. Autoregressive Language Models (ARMs…AppleArbitrage: Efficient Reasoning via Advantage-Aware SpeculationModern Large Language Models achieve impressive reasoning capabilities with long Chain of Thoughts, but they incur substantial computational cost during inference, and this motivates techniques to improve the performance…AppleScaling Categorical Flow MapsContinuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they unlock a host of advantages currently reserved for continuous modalit…AnthropicImproving Fable 5's biology safeguardsGitLabConfidential AI for GitLab Self-HostedYour developers want AI coding agents. Your source code is regulated IP that can t be sent to a third-party AI service, and your compliance team has said so in writing. The usual escape hatch, standing up your own GPU cl…GitLabGitLab Secrets Manager adds ESO, Terraform, API supportToday, you might maintain separate secret stores for CI/CD, Kubernetes, and Terraform. However, that leaves multiple tools to manage, access models to keep in sync, and audit trails to correlate when something goes wrong…GitHubA guide to slash commands in the GitHub Copilot appGo beyond chat in the GitHub Copilot app with these slash commands. They'll help you plan, collaborate, automate, and customize your dev workflow. The post A guide to slash commands in the GitHub Copilot app appeared fir…GitHubHow we took malware advisories beyond npmGitHub malware advisories no longer stop at npm. Here's how we wired OpenSSF's malicious-packages data into the Advisory Database, and why we built the pipeline paranoid. The post How we took malware advisories beyond np…DiscordGeneral Availability of Mobile Platform Support Arrives in Discord Social SDK Version 1.10Mobile support is now generally available in the Discord Social SDK, allowing developers to extend Discord s social experiences to iOS and Android.DiscordDiscord Patch Notes: August 4, 2026Check out the finer details of the more technical fixes implemented into Discord recently.Google DeepMindWeatherNext: AI model achieves breakthrough in forecasting cyclonesCohereCohere and the University of Waterloo launch partnership to strengthen Canada’s AI talent pipelineAWSRuntime instances: persistent compute for production AI agents on Amazon Bedrock AgentCoreAnnouncing runtime instances in Amazon Bedrock AgentCore—persistent, managed EC2 infrastructure for production AI agents with multi-agent collaboration, GPU support, and sessions lasting up to 14 days.AppleDeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer CompletenessLarge language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often struggle to produce complete answer set to complex questions such as “Which actor from…AppleLocking Pretrained Weights via Deep Low-Rank Residual DistillationThe quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and software platforms. They also all…SalesforceHow Salesforce Eliminated Single-Region Risk and Reduced Downtime Blast Radius at 4B Metrics/MinBy Priyanka Singla, Aaron Stockton, and Vincent Poon In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Priyanka Singla, Senior Software…MetaFrom User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads RankingEvery day, Meta’s recommendation platforms handle billions of user interactions, generating rich temporal signals that capture individual preferences and intent across products, ads, and content. In our 2024 post on sequ…AWSAmazon DynamoDB now supports real-time vector search at any scaleDynamoDB now supports native vector search with single-digit millisecond latency at 99%+ recall. It is designed for any scale, even trillions of vectors and requires zero infrastructure management.AppleTaming Outlier Tokens in Diffusion TransformersWe study outlier tokens in Diffusion Transformers (DiTs) for image generation. Prior work has shown that Vision Transformers (ViTs) can produce a small number of high-norm tokens that attract disproportionate attention w…YelpMigrating a Large Flow Monorepo to TypeScriptIn early 2017, Webcore selected Flow as Yelp’s next-generation typechecker over TypeScript. At the time there was no clear frontrunner. Flow had better support for React, better performance, and a respectable repository…Mistral AIIntroducing Shieldstral.GitHubTurn one giant AI-generated pull request to a reviewable stackInstead of one huge, un-reviewable pull request, teach coding agents to decompose work into a clean, ordered stack with GitHub stacked pull requests. The post Turn one giant AI-generated pull request to a reviewable stac…GitHubHow the GitHub legal team used Copilot CLI to streamline their workflowsLearn how to build tools to simplify how you work—without writing a single line of code. The post How the GitHub legal team used Copilot CLI to streamline their workflows appeared first on The GitHub Blog .DiscordSquircles, Styles, and Spacing: How Your Feedback is Helping Improve MobileWe re working to more closely match your mobile and desktop experiences, including changes to server icons and the chat bar. (Did we mention Ash, the OG dark theme, is back?)AnthropicTino Cuéllar joins as Chief Global Affairs OfficerSalesforceRemoving the Security Barrier to Agentforce AdoptionBy Pallavi Rajan Udmalpet and Janhavi Deshpande. As enterprises accelerate secure AI adoption with Agentforce and Data 360, organizations across healthcare, banking, government, and critical infrastructure face a soberin…Microsoft ResearchOrchard: An open framework for scalable agentic AIOrchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to r…MetaGEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation ModelMeta s Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. This post goes i…AWSAWS Weekly Roundup: Price reduction of GPT models in Bedrock, CloudWatch managed collectors for Prometheus metrics, and more (August 3, 2026)Last week I had the joy of participating in Amazon’s “Bring Your Kids to Work Day” with my 7 year old son. We commuted together into the New York City office, his first real rush hour train ride, and spent the day explor…AppleUnderstanding Alignment in Multimodal LLMs: A Comprehensive StudyPreference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively underexplored. Similar to…GrabHow AI is transforming analytics at GrabIntroduction At Grab, analytics sits close to almost every decision that matters. Our north star is the democratisation of intelligence, ensuring that anyone making a business call has immediate access to trustworthy ans…CohereLuke RossNetflixModeling Device Capabilities for Analyticsby Aarti Laddha , Richard Diaz-Cool , Rishika Idnani , Venkatesh Selveraj Netflix supports a vast and evolving set of features and content types, ranging from 4K streaming and immersive audio to live streaming and cloud…GitHubDon&#8217;t stop early: Case-folding source code at memory speedHow a branch-free loop and byte-space arithmetic let GitHub case-fold every byte of code search at >45 GiB/s on a single core. The post Don t stop early: Case-folding source code at memory speed appeared first on The Git…EtsyKafka App? There’s a Skill for ThatEtsy is home to over 100 million listings from 5.6 million active sellers. Because the items for sale are unique and creative, there is no standard product catalog that tells us what they are. When someone searches for “…CohereCohere signs EU Code of Practice on Transparency of AI-Generated ContentSalesforceHow Salesforce Built an Agentic Engineering Enablement Strategy for Thousands of Software EngineersGiving thousands of software engineers access to agentic tooling is straightforward. Helping them fundamentally change how they build software is another matter entirely; one every engineering organization eventually run…NetflixGenRec: Towards LLM-Native Recommendation at NetflixAuthors: Ying Li , Arjun Rao , Shradha Sehgal Introduction Recommendations sit at the heart of the Netflix experience. Our current production models rely on thousands of hand‑crafted features over users, items, and inter…Microsoft ResearchEchoverse: Deep, evolving environments for computer-use agentsComputer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the t…Microsoft ResearchEvoLib: Turning experience into evolving knowledgeLLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib:…GrabCrowdsourced taxonomy verification: A feedback-driven framework for refining knowledge graph relationships via online search interactionsIntroduction The efficacy of semantic search relies on the accuracy of the underlying Knowledge Graph (KG). In high-velocity domains like on-demand food delivery or e-commerce, the catalog of entities like dishes, produc…GitHubStacked sessions and pull requests in the GitHub Copilot appLearn how I modernized an old codebase of mine using stacked sessions and pull requests in the GitHub Copilot app. The post Stacked sessions and pull requests in the GitHub Copilot app appeared first on The GitHub Blog .Google DeepMindGemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaborationGemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.CohereHow AI Is Used in ManufacturingCohereAI in the Healthcare SectorAppleDimensionality Reduction Meets Network Science: Sensemaking on UMAP’s kNN GraphWhile UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally. This…AnthropicInvestigating three real-world incidents in our cybersecurity evaluationsAnthropicInvestigating three real-world incidents in our cybersecurity evaluationsGitHubTame Dependabot: Group your updates, slow the cadence, keep security fastDependabot keeps your dependencies current, but its defaults can flood your repository with pull requests. Here's how grouping updates, slowing the cadence, and keeping security fixes fast cut the noise on a Microsoft op…DiscordIntroducing Profile Frames: Decorative Borders to Make Your Discord Profile Museum-WorthyIntroducing Profile Frames: a new way to put a finishing touch on your Discord profile. Frames add a decorative border around your profile, giving the whole thing a look that s unmistakably yours.Google DeepMindWe’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative controlBAIR (Berkeley)From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple SiliconFigure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied instruction-for-instruction. We face a new epoch in computing.…GitHubDisrupting supply chain attacks on npm and GitHub ActionsExplore the changes we've shipped across npm and GitHub Actions over the past few months to disrupt supply chain attack techniques and limit their impact. The post Disrupting supply chain attacks on npm and GitHub Action…Google DeepMindGemini Robotics 2 brings whole body intelligence to robotsCohereWhat Are AI Agents? How They Work and ExamplesCohereWhat Is Agentic AI? Definition and ExamplesAnthropicDiscovering cryptographic weaknesses with ClaudeAirbnbEval-driven development: Lessons from evaluating GenAI at scaleHow Airbnb teams build trustworthy Generative AI products by treating evaluation as a first-class engineering discipline; not an afterthought. Nestled into the lush hillside, this stunning modern retreat features strikin…SpotifyIndexing the Data Lake for Online Point QueriesCompanies like Spotify need vast quantities of data accessible at low latency for online services and,... The post Indexing the Data Lake for Online Point Queries appeared first on Spotify Engineering .SalesforceBuilding Reliable Production AI with Durable WorkflowsBuilding an AI prototype has never been easier. You send a prompt to a model, receive a response, and present the result to a user. As long as each request is independent, the interaction is straightforward to reason abo…PinterestPinner Progression: Better Use-Case Representation Driving Weekly Active User Growth at PinterestPart 1 of 2 Authors Personalization (Homefeed): Yuke Yan, Chuxi Wang, Andreanne Lemay, Olafur Gudmundsson, Anna Kiyantseva, Krystal Benitez, Jongho Kim, Jiacong He, Rahul Goutam, James Li, Dylan Wang User Understanding:…GitHubThe harness is all you need (mostly)A practical GitHub Copilot workflow for prototyping, planning, implementing, and reviewing software without chasing every new AI tool. The post The harness is all you need (mostly) appeared first on The GitHub Blog .GitHubGitHub Copilot app for Beginners: Getting startedNew to the GitHub Copilot app? Learn how to start projects, work with AI agents, explore canvases, and streamline your development workflow. The post GitHub Copilot app for Beginners: Getting started appeared first on Th…CohereA Day in the Life of a Wealth Manager, With and Without AICohereOrchestrate Agentic Workflows with North AutomationsAWSAWS Weekly Roundup: Local Zone in Athens, Claude Opus 5 on AWS, Lambda durable execution for .NET, and more (July 27, 2026)Last week I had the privilege of spending three days in São Paulo with technical builders from across Latin America, brought together for a regional tech event full of deep-dive sessions, hands-on workshops, and conversa…AnthropicOur position on open-weights modelsAnthropicExpanding our partnership with CognizantAnthropicOur position on open-weights modelsAnthropicOur position on open-weights modelsAnthropicExpanding our partnership with CognizantAnthropicExpanding our partnership with CognizantAnthropicExpanding our partnership with CognizantAnthropicOur position on open-weights modelsBAIR (Berkeley)Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction.abbel-fig { display: block; text-align: center; margin: 2.4em 0; line-height: 1.4; max-width: 100%; } .abbel-fig img { display: block; margin: 0.65em auto 0; height: auto; max-width: 100%; } /* Image sizes; captions use…GrabAgent platform (Part 1): How we help Grab build and run AI agents at scalePart 1: From one support bot to a framework At Grab, AI agents have evolved from interesting team prototypes into production services used every day by millions of merchants, drivers, and consumers. Today, more than 500…AnthropicIntroducing Claude Opus 5AnthropicIntroducing Claude Opus 5AnthropicIntroducing Claude Opus 5AnthropicIntroducing Claude Opus 5SalesforceHow AI Rebuilt Salesforce&#8217;s Decades-Old Localization PipelineIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Teresa Marshall, Vice President of Localization. Teresa s team delivers every Salesfo…MicrosoftThe Microsoft 365 Copilot Agent’s Playbook: A Practical Livestream Series for Building Better AgentsBuilding on Microsoft 365 Copilot? Here s your playbook. Declarative agents are quickly becoming one of the most exciting ways to extend Microsoft 365 Copilot and bring organizational knowledge, workflows, and tools dire…PinterestSecuring Infrastructure at Scale: Introducing Pinterest’s Resource Provisioner Pipeline (RPP)Ammar Ekbote | Senior Software Engineer Chan Kim | Senior Software Engineer Managing Infrastructure as Code (IaC) across a massive organization comes with a unique set of security and logistical challenges, particularly…Google DeepMindAccelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis MissionGoogle commits $40M in AI tokens and credits for the Genesis MissionAnthropicThe Anthropic Economic Index connectorAnthropicA research agenda for the Economic Futures Research FundAnthropicA research agenda for the Economic Futures Research FundAnthropicThe Anthropic Economic Index connectorAnthropicA research agenda for the Economic Futures Research FundAnthropicA research agenda for the Economic Futures Research FundAnthropicThe Anthropic Economic Index connectorStripeAnalyzing the evidence that helps businesses win “product not received” disputesTo understand what can influence win rates, we analyzed evidence packets from one million disputes over a 16-week period. Here’s what the data shows and what it means for how you mitigate disputes.SalesforceHow AI Reduced Customer Bug Triage from Nearly a Year to Less Than a WeekBy Priya Sethuraman, Abhishek Ghose, Lovish Agarwal, and Aditya Pandey. In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Priya Sethura…MicrosoftHow to test agent experience changes without shipping themMost changes you think will improve AI agent behavior won't. We tested a dozen hypotheses on a real project upgrade scenario and the majority failed. Learn how to emulate documentation, API, and MCP server changes locall…Google DeepMindIntroducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash CyberWe’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.Google DeepMindIntroducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash CyberWe’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.AnthropicDonating another $20 million to Public First ActionAnthropicDonating another $20 million to Public First ActionAnthropicDonating another $20 million to Public First ActionAirbnbPersonalizing Airbnb search by learning from the guest journeyHow we built a Transformer-based sequence model that encodes years of guest behavior to surface the right listings at the right time. By: Daochen Zha , Chun How Tan , Xin Liu , Bin Xu , Han Zhao , Xiaowei Liu , Jun Shi ,…SpotifyContent Ingestion & Podcast Video Incident ReportOver the past two months, podcast creators have experienced a series of reliability issues on Spotify. This... The post Content Ingestion Podcast Video Incident Report appeared first on Spotify Engineering .DropboxHow our universal content processing platform Riviera evolved for AI and beyondRiviera is the Dropbox content processing platform that’s been iteratively improving content transformation in our products for roughly a decade.AWSAWS Weekly Roundup: One-click Lambda setup prompt, OpenAI GPT-5.6 models on Bedrock, and more (July 20, 2026)Last week, my team visited Seoul to meet AWS Korea User Group (AWSKRUG) leaders. AWSKRUG is the largest cloud developer community in Korea, with 20 meetup groups organized by topic and area that collectively host over 10…SalesforceClosing the Loop: How to Build Self-Improving AI Systems with Automated Feedback LoopsAfter six cycles, the system ran out of things to fix. That sentence probably raises more questions than it answers. Here is how we got there. The first skill PR took three days to merge. Twelve comments on structure. Si…NetflixIn-House LLM Serving at NetflixBy AI Platform’s Model Runtime team and Inference team Introduction Most organizations consume LLMs through hosted APIs. Netflix went further — we run the full stack ourselves, from model deployment through inference, in…MicrosoftHow to test agent skills without hitting real APIsYour agent skill calls an API. The moment you start evaluating it, every run either costs money or mutates production data. Learn how to mock APIs transparently so you can run evals without changing your skill or hitting…Google DeepMindIntroducing Gemini 3.5 Flash CyberGoogle introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.YelpMigrating from Apollo Tooling to GraphQL Codegen at YelpIntroduction At Yelp, we rely heavily on GraphQL and Apollo for data loading in our frontend React monorepo. When a developer writes a GraphQL query or mutation inside a React component, the shape of the response is defi…SalesforceUsing Claude to Build an AI Knowledge Base in 30 MinutesEver had an experience onboarding someone, where you had to provide them a handful of documents, all authored at different times, and in various states of completeness? Ever had your AI agent burn cycles pulling down tha…Google DeepMindOur approach to bioresilienceGoogle DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience and AI models.CohereCohere and the University of Toronto partner to advance responsible AI adoption at scaleMistral AIRobostral Navigate: single-camera AI navigationMicrosoftBuilding AX evals that actually workThis is the eighth and final article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent sta…MicrosoftBuilding AX evals that actually workThis is the eighth and final article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent sta…MetaExploring Hierarchical Interest Representation For Meta Ads Deep Funnel OptimizationHierarchical Interest Representation is a research area for Meta Ads. We’re exploring an upstream representation layer over the universe of Ads entities – users, advertisers, products, services – learning unified embeddi…CohereThe Total Cost of AI Ownership (AI TCO)CohereLanguage, Decoded: What if you could code in your own language?CohereAriana MilliganYelpTraining Orchestrator: Unifying Model Training at YelpAt Yelp, we train many machine learning models on different schedules. Applied machine learning teams all have their own set of Spark-based training batches, scripts, and configurations. Over time, these diverged, leadin…SlackShipyard: How We Built Slack’s Next-Generation EC2 PlatformOver the past few years, we’ve been on a journey to modernise how we run Amazon Elastic Compute Cloud (EC2) instances at Slack. In our first post, Advancing Our Chef Infrastructure, we shared how we moved from a single C…InstacartBlueberry: Force Multiplier For The On-Call EngineerHow we built a Slack-native on-call reasoning harness at Instacart that shortens time to first insight, speeds up theory testing, and turns tribal knowledge into reusable infrastructure. Key Contributors: Karthik Halukur…CohereTiny Aya Expedition Drives Multilingual InnovationCohereFrom Reviewing Words to Building Africa's AI: Our Journey Inside the Cohere Labs Open Science CommunityAirbnbFrom weeks to a day: how we made LLM evaluation fast enough to iterate onTraining an LLM is the easy part. The hard part is designing experiments and evaluations that you can trust enough to know whether the new model is actually an improvement. By : Baharak Saberidokht Introduction Shipping…NetflixBuilding Service Topology at Scale: Architecture, Challenges, and Lessons LearnedBy Parth Jain , Rakesh Sukumar , Yingwu Zhao , Renzo Sanchez-Silva Nathan Fisher A deep dive into the engineering challenges of building a real-time service dependency map at Netflix scale: from streaming architectures a…Microsoft ResearchVerifying Rust cryptography in SymCrypt, from standards to codeCryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The p…MetaModernizing the Meta Ads Service With an Open-Source Kernel SchedulerTL; DR At Meta s scale, a few milliseconds of latency degradation can have a significant negative impact on ads performance. When a Linux kernel upgrade risked regressing latency across Meta s ad serving fleet, we turned…Google DeepMindEmpowering India’s next generation of innovators with ATL SaathiGoogle and AIM launched ATL Saathi, a Gemini-powered AI tool empowering Indian educators in robotics labs.AWSAWS Weekly Roundup: AWS Builder Center at 1 year, Network Scanning in Security Hub, Loom for AWS, and more (July 13, 2026)AWS Builder Center turned one year old last week. Launched on July 9, 2025, the platform has grown from a community hub with Wishlist voting, community profiles, and a toolbox into a full ecosystem with sandbox environme…AWSAmazon SQS turns 20: Two decades of reliable messaging at scaleOn July 13, 2006, we launched Amazon Simple Queue Service (Amazon SQS) as one of the first three services available to customers, alongside Amazon EC2 and Amazon S3. We had learned firsthand that distributed systems need…GrabScaling Grab's Data Lake: Our journey to Apache Iceberg adoptionIntroduction: The evolution of Grab’s Data Lake At Grab’s scale, managing petabytes of data across billions of S3 objects demands more than a storage layer. It demands a robust architectural primitive that supports the h…CohereMulti-Agent Systems: Enterprise AI GuideCohereHardware-Aware, Dynamic Speculative Decoding (DSD)SalesforceHow Informatica Reduced Data Integration Pipeline Development from Days to MinutesIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Nancy Chen, Vice President of Engineering at Informatica. Nancy leads the development…Microsoft ResearchAurora 1.5: Extending open foundation models for weather and Earth-system applicationsAurora 1.5 adds 22 more variables, hourly temporal resolution, and probabilistic ensemble forecasting to the Aurora foundation model, making it more useful for real-world weather, climate, and energy applications. The po…Mistral AIVersion control for prompts & skills in StudioLyftFrom Day 1 to Production: Building Lyft’s Analytics & Rides Intelligence Assistant as Onboarding…From Day 1 to Production: Building Lyft’s Analytics Rides Intelligence Assistant as Onboarding Project Written by Sagar Baronia at Lyft. A Different Kind of Day One Most onboarding journeys follow a familiar arc: orienta…Microsoft ResearchFlint: A visualization language for the AI eraShort chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, huma…MicrosoftLet’s Learn GitHub Copilot App &#8211; Free Virtual Training EventJoin us for a free online event series kicking off July 16 to learn how to get started with the GitHub Copilot App! The post Let’s Learn GitHub Copilot App Free Virtual Training Event appeared first on Microsoft for Deve…MicrosoftThe hidden variables in your agent evalThis is the seventh article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how t…MicrosoftLet’s Learn GitHub Copilot App &#8211; Free Virtual Training EventJoin us for a free online event series kicking off July 16 to learn how to get started with the GitHub Copilot App! The post Let’s Learn GitHub Copilot App Free Virtual Training Event appeared first on Microsoft for Deve…MicrosoftThe hidden variables in your agent evalThis is the seventh article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how t…CohereHow to build an AI agent: A practical guide for enterprisesMicrosoftDon&#8217;t rewrite your CLI for agentsThere s advice making the rounds: replace your CLI args with a single --json payload so agents can use your tool more effectively. The thinking being, that agents already think in structured formats, and nested data maps…MicrosoftDon&#8217;t rewrite your CLI for agentsThere s advice making the rounds: replace your CLI args with a single --json payload so agents can use your tool more effectively. The thinking being, that agents already think in structured formats, and nested data maps…DiscordDiscord Patch Notes: July 7, 2026Check out the finer details of the more technical fixes implemented into Discord recently.CohereCohere Transcribe Arabic: Open-Source Speech AICohereInside Tiny Aya: What an SAE Sees in a Model Built for 70+ LanguagesBAIR (Berkeley)Intelligence is Free, Now What? <br> Data Systems for, of, and by Agents... government of the people, by the people, for the people ... Abraham Lincoln, Gettysburg Address (1863) The cost of AI is dropping rapidly. GPT-4-class capabilities cost roughly $30 per million tokens in early 2023; t…SalesforceBuilding Enterprise AI Agents That Are Both Autonomous and ReliableWhile red-teaming an early refund agent early last year, one of our engineers typed a deliberately absurd line: my very valid and very verified email. The agent, running on a frontier LLM, accepted it as proof of identit…MicrosoftNot all model upgrades are upgradesA new model drops with lower per-token pricing and better benchmarks. You switch. A week later someone asks why the agent is burning 12x more tokens on the same task while producing worse output. We ran 150 agent tasks a…MicrosoftNot all model upgrades are upgradesA new model drops with lower per-token pricing and better benchmarks. You switch. A week later someone asks why the agent is burning 12x more tokens on the same task while producing worse output. We ran 150 agent tasks a…AWSAWS Weekly Roundup: Claude Sonnet 5 on AWS, Amazon WorkSpaces for AI agents, AWS service availability updates, and more (July 6, 2026)A couple of editions ago I wrote about what I find so energizing about working with startups. Last week I got a fresh dose of it: I spent a few days with the AWS Startups team, listening to stories of founders talking ab…GrabMigrating Counter Service storage: Design choices and learningsIntroduction Counter Service is used across Grab’s anti-fraud platform to answer time-windowed count questions, such as recent ride requests by a user or failed payment attempts on a card. The service handles tens of tho…Google DeepMindGoogle DeepMind and A24 announce first-of-its-kind research partnershipCohereSinhala Is Not Just Low-Resource: It Is Under-EvaluatedMistral AILeanstral 1.5: Proof Abundance for AllSalesforceHow AI Learned to Investigate Mobile Build Failures Like an Experienced Support EngineerIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Archana Indran, Senior Engineering Manager for Mobile CI/CD. Archana s team built Ana…MicrosoftWhat AI benchmarks are not telling youThis is the sixth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how to…MicrosoftWhat AI benchmarks are not telling youThis is the sixth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how to…MetaMeta&#8217;s AI Storage Blueprint at ScaleOver the past several years, model capabilities and training dataset sizes have experienced exponential growth. During the past year or so, the time between new-frontier-model releases has gone down from months to weeks.…InstacartVariance Reduction Below the Randomization GrainSergio Camelo, Caitlin Kearns, Matias Cersosimo, and Tilman Drerup As artificial intelligence increases the velocity of engineering and science teams, experimental throughput is set to become a bottleneck for many produc…BAIR (Berkeley)2026 BAIR Graduate ShowcaseCongratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026! This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed…AWSUpgrade Amazon EKS clusters with confidence using Kubernetes version rollbacksLearn how Kubernetes version rollbacks for Amazon EKS let you reverse cluster upgrades within seven days. This new feature provides a safety net for upgrade failures—no cluster rebuilds required—turning Kubernetes versio…Microsoft ResearchSkillOpt: Agent skills as trainable parametersAI agents often fail because their instructions, or skills, are manually modified with no guarantee of improvement. Learn how SkillOpt turns skill editing into a training process, making agent behavior more reliable with…Meta10 Years of Meta’s Commitment to PythonThis year marks Meta s 10th consecutive year as a sponsor of the Python Software Foundation (PSF), the charitable organization dedicated to advancing, supporting, and protecting the open-source Python programming languag…DiscordCost Attribution in Discord’s APIDiscord s API spans 1700+ endpoints across hundreds of Kubernetes deployments. The challenge: tracking per-feature hosting costs without restructuring. Jim Benton helps explain how Discord tackled the situation.DiscordDiscord is Now on Meta Quest: Reach Out to Your Servers While in VRStarting today, Discord is now available to download directly from Meta’s Horizon App Store. No more sideloading. No more using the web app. Grab it using the link in this blog, or at discord.com/download.Google DeepMindStart building with Nano Banana 2 Lite and Gemini Omni FlashAWSAccelerate your infrastructure deployments by up to 4x with AWS CloudFormation Express modeAWS CloudFormation speeds up infrastructure deployment with Express mode, enabling AI agents and developers to receive deployment confirmation in seconds and iterate faster. Available in all commercial Regions at no addi…AWSAmazon EC2 C9g and C9gd instances powered by AWS Graviton5 processors are now availableAmazon EC2 C9g and C9gd instances, powered by AWS Graviton5, are now generally available. They deliver up to 25% better compute performance than Graviton4-based instances, 5x larger cache, fastest memory of any processor…AWSAutomate public TLS certificate issuance with ACME support in AWS Certificate ManagerAWS Certificate Manager now supports the ACME protocol for public TLS certificates, enabling automated issuance and renewal through any ACMEv2-compatible client on any workload. Administrators get centralized governance,…SalesforceInside Unified Planner: The AI Brain Behind AgentforceIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Gaurav Aggarwal, a Software Engineering Architect. Gaurav developed the Unified Plann…NetflixGenPage: Towards End-to-End Generative Homepage Construction at NetflixAuthors: Lequn Wang , J iangwei Pan , and Linas Baltrunas Figure 1. Autoregressive homepage generation. GenPage builds a Netflix homepage one row or entity at a time, each one conditioned on what’s already on the page an…Microsoft ResearchMemora: A Harmonic Memory Representation Balancing Abstraction and SpecificityAI agents can't remember past conversations. They must constantly reload or retrieve context, which grows less efficient as tasks get longer and more complex. Memora solves this with a scalable memory system separating w…InstacartLeveraging PyFixest for High-Cardinality Marketplace Modeling at InstacartBenjamin S. Knight Scaling Marketplace experiments requires specialized statistical techniques. We examine why standard ordinary least squares regression (OLS) becomes computationally intractable when controlling for hig…AWSAWS Weekly Roundup: Agentic CX designer for Amazon Connect Customer, EC2 AMI Watermarks, Open Governance for MySQL, and more (June 29, 2026)It has been a busy stretch on the AWS Summit circuit. At the New York City Summit, I delivered a workshop called Building AI architectures with AWS Serverless, and it was a lot of fun watching builders wire up agents and…MicrosoftYour agent already has a planIf an agent isn t doing the right thing, the obvious move is to make the docs clearer. Add a tip, spell out the correct command, describe the right approach more prominently. You do all of that, and the agent still ignor…MicrosoftYour agent already has a planIf an agent isn t doing the right thing, the obvious move is to make the docs clearer. Add a tip, spell out the correct command, describe the right approach more prominently. You do all of that, and the agent still ignor…PinterestAchieving Near-Linear Training Scalability for Pinterest’s Foundation ModelsSheng Huang | Software Engineer, AI Platform; Pong Eksombatchai | Machine Learning Engineer, Applied Sciences; Saurabh Vishwas Joshi | Software Engineer, AI Platform; Gaurav Arora | Software Engineer, AI Platform; Karthi…Microsoft ResearchUnderstanding the brain with AI-driven explanations and experimentsResearchers introduce generative causal testing, which translates black box models into clear hypotheses and verifies them in the scanner, revealing what specific brain regions respond to in language. The post Understand…MicrosoftLearn from Microsoft: Transform software development through an agentic platformSee how Microsoft is transforming software development with agentic workflows, AI-powered automation, and specialized agents across the engineering lifecycle. The post Learn from Microsoft: Transform software development…MicrosoftLearn from Microsoft: Transform software development through an agentic platformSee how Microsoft is transforming software development with agentic workflows, AI-powered automation, and specialized agents across the engineering lifecycle. The post Learn from Microsoft: Transform software development…MetaPrivacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case StudyPrivacy controls — systems that enforce retention, access, allowed-purpose, downstream-sharing, or anonymization policies — require a reliable understanding of data to function. Before such a control can operate effectiv…DropboxHow we used DSPy to turn AI evaluations into better responses in Dash chatWe used DSPy to improve LLM judges and optimize our chat experience, creating an evaluation-driven feedback loop that produced better outputs.DiscordDiscord Update: June 25, 2026 ChangelogHere s the Discord Changelog from June 25, 2026, so you can stay informed on what’s new in recent app updates!CohereAutomating Fork Maintenance with AI AgentsCohereCreating a Security Agent with Cohere North and WizPinterestAutomated Schema Evolution in Pinterest’s Next-Generation DB Ingestion FrameworkYisheng Zhou | Software Engineer II Liang Mou | Sr Staff Software Engineer Gabriel Raphael Garcia Montoya | Staff Software Engineer Istvan Podor | Staff Software Engineer Introduction In the first post of this series , w…Microsoft ResearchTalos: Scaling rare disease diagnosis with automated, iterative genomic reanalysisTalos was built to help resolve a major bottleneck in genomic medicine: human review time. The open-source system recovered 90% of in-scope diagnoses while surfacing just 1.3 candidate variants per patient for expert rev…Mistral AIBringing more control over your connectorsMicrosoftWhen the model has never seen your codeThis is the fifth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how to…MicrosoftWhen the model has never seen your codeThis is the fifth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how to…Google DeepMindIntroducing computer use in Gemini 3.5 FlashCohereWhen AI Doctors "See" What Isn't There: Why Better Accuracy Doesn't Mean Better VisionCohereAnanya SahuStripeFour travel and hospitality trends from HITEC 2026More than 6,000 hospitality executives and operators gathered in San Antonio last week for the HITEC conference. The big topic: whether the industry’s AI investment is actually working. Across four days and over 50 meeti…SalesforceHow Agentforce Prevents Language Drift in 600K Daily Multilingual AI WorkflowsIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Ishween Kaur, Senior Software Engineer on the Agentforce Agentic Reasoning team. As A…NetflixToward More Controllable AI Video Editing: An Early Research Exploration at NetflixBy Zhuoning Yuan , Ta-Ying Cheng , Benjamin Klein , Bahareh Azarnoush Introduction At Netflix, we build technology to help storytellers bring their creative visions to life and to help members discover the stories they l…MozillaPACT: Anonymous Credentials for the WebThis is the technical companion to our update on Distilled, “Keeping the web open and private in the bot era.” Here we take a deeper look at the problem space, the design we re proposing, and the problems still left to s…Mistral AIMistral OCR 4 : SOTA OCR for Document IntelligenceMistral AIOCR 4MetaHow Meta Engineered Ultra-Narrow Batteries for AI GlassesSmart glasses like the Ray-Ban Meta and Oakley Meta Vanguards need to pack enough energy to power features like cameras, speakers, AI workloads, and even a display. But it all has to fit into the glasses’ temple arms. So…NetflixHow Netflix Simplified Batch Compute with KueueBy Alvin Bao , Alex Petrov , Jennifer Lai , Aidan Sherr , and Samartha Chandrashekar As a part of the journey to transition Netflix’s compute infrastructure to be more Kubernetes-native, we have leaned into incorporating…MicrosoftModels don&#8217;t have preferences, they have contextYou open a fresh chat, type What framework should I use for a web app? , and the model says React. You screenshot it, share it, and write Claude prefers React. It gets engagement. People nod along. A few reply with their…MicrosoftModels don&#8217;t have preferences, they have contextYou open a fresh chat, type What framework should I use for a web app? , and the model says React. You screenshot it, share it, and write Claude prefers React. It gets engagement. People nod along. A few reply with their…MetaAdopting AV1 for Real-Time Communication (RTC) at ScaleAdopting AV1 for real-time communication at Meta has been a multi-year effort spanning codec selection, device eligibility, rate control, and error resilience. We’re sharing the technical and operational challenges while…GrabScaling out Distroless adoption With AIDistroless adoption at Grab Grab is migrating from heavy base images to Distroless images to reduce security risks. By limiting each container to the application and its runtime dependencies, we shed non-essential binari…AWSRun isolated sandboxes with full lifecycle control: AWS Lambda introduces MicroVMsAWS launches a new serverless compute primitive, AWS Lambda MicroVMs. VM-level, isolated sandboxes with no shared kernel or resources between sessions. Rapid launch and resume, full lifecycle control, state preservation…AWSAWS Weekly Roundup: NY Summit recap, Local Zone in Hanoi, Grok 4.3 in Bedrock, price reductions, and more (June 22, 2026)Last week AWS Summit New York City brought together thousands of customers, partners, and builders for a free, one-day event showcasing the latest in cloud and AI innovation. Dr. Swami Sivasubramanian, VP of Agentic AI a…GrabPalana (Part 2): Architecting isolation, identity, and auditability for AI agentsIntroduction In Part 1 , we introduced Palana , Grab’s Kubernetes-native secure execution platform for autonomous AI agents. We discussed the underlying need for isolated environments and covered its core design principl…NetflixThe Data Canary: How Netflix Validates Catalog MetadataBy Celina Amados At Netflix, our catalog metadata is crucial to our member experience, and a single corrupted data state can impact millions of viewers immediately. To protect streaming reliability, we built an automated…NetflixData Projects: Managing Data Assets at Netflix ScaleBy Amer Hesson , Marcelo Mayworm , James Mulcahy , and Brittany Truong The Problem: Managing Assets at Netflix Scale Netflix’s Data Platform is vast. We have millions of tables in our data warehouse and tens of thousands…NetflixPredicting Risk in Content Launches: How Data-Driven Insights can Transform Launch Planningby Emily Gill Each year, we bring the Analytics Engineering community together for an Analytics Summit — a multi-day internal conference to share analytical deliverables across Netflix, discuss analytic practice, and bui…NetflixThinking Fast & Slow for a Personalized Notification Systemby Matthew Wood , Ishan Gupta , Kevin Mercurio, Devon Bryant , and Claire Dorman In his seminal book “Thinking, Fast and Slow,” Daniel Kahneman describes two systems that drive human cognition: System 1, which operates a…NetflixVMAF v1: Good Is Not Good EnoughBy Christos G. Bampis , Zhi Li , Kyle Swanson, Nil Fons Miret and Pavan Madhusudanarao Will this encode look good to Netflix members? Does switching to a new codec improve quality at the same bitrate and by how much? Wha…NetflixFrom Silos to Service Topology: Why Netflix Built a Real-Time Service MapBy Parth Jain , Rakesh Sukumar , Yingwu Zhao , Renzo Sanchez Nathan Fisher How we built a living map of our distributed infrastructure to help engineers understand dependencies, troubleshoot faster, and keep Netflix runn…NetflixA Human-Augmenting Agentic Workflow for Causal InferenceBy Winston Chou, Adrien Alexandre, Lars Olds, Yi Zhang, Garrett Hagemann, and Nathan Kallus Introduction Imagine asking a data agent to analyze the causal relationship between two variables, such as the effect of watchin…NetflixThinking Fast & Slow for a Personalized Notification Systemby Matthew Wood , Ishan Gupta , Kevin Mercurio, Devon Bryant , and Claire Dorman In his seminal book “Thinking, Fast and Slow,” Daniel Kahneman describes two systems that drive human cognition: System 1, which operates a…NetflixThe Evolution of Cassandra Data Movement at NetflixBy Guil Pires , Jennifer Prince , Jose Camacho , Ken Kurzweil , Phanindra Chunduru Background In a previous post, we introduced Data Bridge , a unified management plane for batch Data Movement at Netflix. Historically, s…NetflixPredicting Risk in Content Launches: How Data-Driven Insights can Transform Launch Planningby Emily Gill Each year, we bring the Analytics Engineering community together for an Analytics Summit — a multi-day internal conference to share analytical deliverables across Netflix, discuss analytic practice, and bui…NetflixData Projects: Managing Data Assets at Netflix ScaleBy Amer Hesson , Marcelo Mayworm , James Mulcahy , and Brittany Truong The Problem: Managing Assets at Netflix Scale Netflix’s Data Platform is vast. We have millions of tables in our data warehouse and tens of thousands…NetflixThe Data Canary: How Netflix Validates Catalog MetadataBy Celina Amados At Netflix, our catalog metadata is crucial to our member experience, and a single corrupted data state can impact millions of viewers immediately. To protect streaming reliability, we built an automated…NetflixVMAF v1: Good Is Not Good EnoughBy Christos G. Bampis , Zhi Li , Kyle Swanson, Nil Fons Miret and Pavan Madhusudanarao Will this encode look good to Netflix members? Does switching to a new codec improve quality at the same bitrate and by how much? Wha…NetflixFrom Silos to Service Topology: Why Netflix Built a Real-Time Service MapBy Parth Jain , Rakesh Sukumar , Yingwu Zhao , Renzo Sanchez-Silva Nathan Fisher How we built a living map of our distributed infrastructure to help engineers understand dependencies, troubleshoot faster, and keep Netfli…NetflixA Human-Augmenting Agentic Workflow for Causal InferenceBy Winston Chou, Adrien Alexandre, Lars Olds, Yi Zhang, Garrett Hagemann, and Nathan Kallus Introduction Imagine asking a data agent to analyze the causal relationship between two variables, such as the effect of watchin…NetflixThe Evolution of Cassandra Data Movement at NetflixBy Guil Pires , Jennifer Prince , Jose Camacho , Ken Kurzweil , Phanindra Chunduru Background In a previous post, we introduced Data Bridge , a unified management plane for batch Data Movement at Netflix. Historically, s…GrabPalana (Part 1): Why Grab built a secure platform for autonomous AI AgentsAbstract Artificial intelligence (AI) agents are moving from experiments into everyday engineering workflows. They can read code, call application programming interfaces (APIs), run tests, create merge requests, answer S…StripeWhat Link data tells us about AI spendingWe analyzed spending patterns across the 250 million customers paying with Link. We found that Link customers are spending more on AI than they were three months prior, investing heavily in platforms that let them build…Mistral AIFrontier AI LLMs, assistants, agents, servicesMicrosoftStop overloading your skillsYou built a skill for your technology. API references, authentication flows, SDK patterns, error handling, version info, all packed into one skill. The agent calls it, gets all that context, and generates code. The kicke…DiscordHow to Manage Your Discord Desktop Notifications: A Complete GuideNot getting pinged for the conversations you wanna know about? This guide walks through server and channel notifications settings, Do Not Disturb, turning off specific notification sounds, and what to check when your pin…AWSAmazon ECS introduces new high-resolution metrics for faster service auto scalingAmazon Elastic Container Service (Amazon ECS) service auto scaling automatically adjusts task counts to meet workload demand with comprehensive scaling policies, including predictive scaling for recurring traffic pattern…AWSAnnouncing Amazon EC2 G7 instances accelerated by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUsAnnouncing the general availability of Amazon Elastic Compute Cloud (Amazon EC2) G7 instances, delivering high performance GPU acceleration for AI inference, graphics, and data analytics workloads.SalesforceMaintaining Code Quality at Agent Speed: 7 Patterns for Agentic EngineeringBy Amit Sharma and Antonio Garrote.How do you know whether code generated at agent speed can be trusted? That question is fast becoming one of the most important in software engineering and it deserves a serious answer.…MicrosoftWhen your agent extensions fight each otherThis is the fourth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how to…AWSIntroducing Amazon Bedrock Managed Knowledge Base for faster, more accurate enterprise AI applicationsAmazon Bedrock's new Fully Managed Knowledge Bases simplifies building enterprise RAG pipelines by providing native data connectors Smart Parsing for automatic multi-format data preparation, and an Agentic Retriever for…AWSAWS DevOps Agent adds release management capabilities to assess code changes before production (preview)AWS DevOps Agent now offers release management capability in preview, reviewing code changes for release readiness and running autonomous release testing to help you ship code to production safely and with confidence.AWSAWS Security Agent adds threat modeling, Kiro power and Claude Code plugin, and moreAWS Security Agent now adds STRIDE-based threat modeling, full repo and PR code scanning with remediation across major Git platforms, and IDE integrations via Kiro power, Claude Code plugin, and MCP — letting developers…AWSProactively reduce tech debt autonomously with AWS Transform – continuous modernization (preview)AWS Transform – continuous modernization (preview) automatically scans code repositories to detect, prioritize, and remediate technical debt at scale.AWSAnnouncing Web Search on Amazon Bedrock AgentCore: Ground your AI agents in current, accurate web knowledgeAWS introduces Web Search on Amazon Bedrock AgentCore, a fully managed tool that enables agents to ground responses in current, cited web knowledge with zero data egress from customer's secured AWS environment. You can f…AWSTop announcements of the AWS Summit in New York, 2026A recap of the top announcements from AWS's New York Summit 2026SalesforceThe Agent Coding Maturity Curve: 9 Stages from Code Generation to Trusted AutomationMost developers begin with the same rush of excitement: the agent writes code, fixes bugs, explains unfamiliar systems, generates tests, and turns vague intent into something that looks runnable. For a moment, it feels l…MicrosoftCompeting against yourselfYou shipped a new CLI: better developer experience, modern architecture, and optimized for agents. You deprecated the old one, updated the docs, and blogged about it. Developers are migrating. Then someone asks an AI cod…Google DeepMindSecuring the future of AI agentsSecuring internal systems with an AI Control Roadmap, combining traditional safeguards and real-time monitoring.Google DeepMindUnlocking UK house-building with AI-accelerated planningUK government partners with Google DeepMind to build a new AI-powered prototype aimed at faster housing decisions.AWSAmazon S3 annotations: attach rich, queryable context directly to your objectsAmazon S3 now lets you attach up to 1 GB of rich, mutable, and queryable context directly to your objects using annotations, purpose-built for AI agents and autonomous workflows that need to discover, understand, and act…SalesforceHow Data 360 Segmentation Processes a Quadrillion Records Across Arbitrary Customer Data ModelsIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Deepak Pushpakar, Software Engineering Architect for Segmentation and Activation with…PayPalWe have moved to https://developer.paypal.com/community/blogThe PayPal Developer Blog has a new home. You’ll continue to find the latest developer news, tutorials, product updates, and engineering insights here: https://developer.paypal.com/community/blog Thanks for being part of…Microsoft ResearchIre identifies another LOTUSLITE specimenProject Ire examined a timely malware sample and determined its intent through reverse engineering—identifying LOTUSLITE characteristics even as most major EDR tools did not detect it. The post Ire identifies another LOT…Mistral AIFrontier AI LLMs, assistants, agents, servicesDropboxHow Dropbox uses MCP and Dash to close the design-to-code security gapUsing an agentic AI system to surface threat models during code review and spot gaps between security requirements and implementation.StripeStripe Projects adds new agent integrations, more providers, and custom developer controlsOur data shows that agents are now fully capable of independently writing code and integrating with APIs like Stripe’s. And yet, many of the steps adjacent to writing code are still too hard for agents to do on their own…SlackAgentic Testing: Where Agents Fit in the E2E Testing StackAbstract Agent-driven end-to-end (E2E) tests add a new exploratory layer to testing, but should they replace traditional deterministic tests? We ran more than 200 agentic E2E workflows using the Playwright MCP, Playwrigh…SalesforceHow MuleSoft Is Raising the Trust Bar for AI-Generated CodeIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Melissa Cazalet, Senior Vice President of Software Engineering at MuleSoft, whose tea…MicrosoftYour agent just scaffolded a project from 2020Your agent ran a scaffold command. Project generated, dependencies resolved, no errors. Everything looks fine. Except it s based on the project structure from 2020, and neither you nor the agent noticed. How npx picks th…SpotifyEncoding Your Domain Expert: The Context Layer Behind Spotify's Data AssistantAt Spotify, data problems used to follow a specific pattern. You'd look for the relevant dashboard, there... The post Encoding Your Domain Expert: The Context Layer Behind Spotify's Data Assistant appeared first on Spoti…MicrosoftSpec-Driven Development: A Spec-First Approach to AI-Native EngineeringAI has made software delivery faster, but speed alone does not guarantee better outcomes. As teams adopt AI-native development, the real challenge is keeping requirements, design, implementation, and validation aligned s…MicrosoftStop skillmaxxing, save your tokensYou built a dozen skills for your technology: authentication, CRUD, error handling, deployment, testing, monitoring. Then you installed a cloud platform bundle with 15 more covering diagnostics, storage, compliance, and…MicrosoftIs your agent extension actually working?This is the third article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how to…LyftMetric Semantic Layer: How Lyft Governs and Scales Key Data DefinitionsWritten by Rohit Channe and Simran Mirchandani at Lyft. Motivation At Lyft, data isn’t just a resource — it’s woven into everything we do. Metrics drive key forecasts, steer operational decisions, and put our boldest hyp…LyftMetric Semantic Layer: How Lyft Governs and Scales Key Data DefinitionsWritten by Rohit Channe and Simran Mirchandani at Lyft. Motivation At Lyft, data isn’t just a resource — it’s woven into everything we do. Metrics drive key forecasts, steer operational decisions, and put our boldest hyp…DiscordGame On: Discord Is Backing the Next Generation of Dutch Gaming FoundersTogether with Techleap, we’re launching the Gaming Founders Circle: a hands-on program for the most ambitious gaming companies in the Netherlands to provide participants with a direct line to the people, investors, and n…DiscordUpdated Requirements to How Apps Access Data in ServersDiscord is updating requirements to how apps access certain data within servers, changing the review threshold and requiring annual review. Here s what s changing and why.Google DeepMindInvesting in multi-agent AI safety researchGoogle DeepMind and partners announce a $10M funding call for multi-agent safety research.Google DeepMindDiffusionGemma: 4x faster text generationSalesforceHow to Build Reliable AI Agents: 5 Engineering Patterns from a Production SystemBy Tuhin Kanti Sharma and Chirag Ramesh Hegde. If you ve built an AI agent that works perfectly in demos but becomes unpredictable in production, you ve probably already discovered that reliability is much harder than ca…LyftFrom Chaos to Clarity: How We Built a Unified, Self-Routing Support Ops Ticketing System at LyftWritten by Atul Gupta , Analytics Manager — LUS Support Ops, Lyft At Lyft, getting operators and riders connected quickly and reliably depends on more than technology — it depends on the teams working behind the scenes t…LyftFrom Chaos to Clarity: How We Built a Unified, Self-Routing Support Ops Ticketing System at LyftWritten by Atul Gupta , Analytics Manager — LUS Support Ops, Lyft At Lyft, getting operators and riders connected quickly and reliably depends on more than technology — it depends on the teams working behind the scenes t…DiscordHow We Moved Discord Voice to the EdgeMoving Discord’s voice and video onto Cloudflare s edge network. Closer servers, lower ping in most regions, and a few real bugs getting there.Google DeepMindPowering the future of robotics in EuropeGoogle DeepMindIntroducing Gemma 4 12B: a unified, encoder-free multimodal modelGoogle DeepMindFluid, natural voice translation with Gemini 3.5 Live TranslateGemini 3.5 Live Translate brings near real-time, natural speech translation to Google AI Studio, Google Translate and Google Meet.AirbnbScaling beyond one: How Airbnb evolved its data architecture for a multi-product worldHow Airbnb’s data engineers and analytics engineers built a consistent and flexible data modeling framework to support the expansion into Homes, Experiences, and Services. By : Patrick Lam , Namrata Lamba , Jamie Stober…SalesforceScaling Zero Copy from 1 Trillion to 120 Trillion Rows with File FederationIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Srini Krishnamoorthy, Vice President of Engineering for Data 360. Srini leads the evo…MicrosoftMicrosoft Build 2026 recap: vision, launches, and top sessionsCatch up on Microsoft Build 2026 with the vision lead-off, top developer announcements, and must-watch sessions across the Microsoft developer ecosystem. The post Microsoft Build 2026 recap: vision, launches, and top ses…DiscordIntroducing: You BarIntroducing the You Bar: an update to the Discord mobile app that celebrates your identity, simpler navigation, and a peek at what’s next.Google DeepMindMeasuring the impact of learning with AI in Sierra Leone and beyondResults from a randomized controlled trial show the potential of Gemini’s Guided Learning feature to boost engagement and accelerate learning.StripeRethinking risk in the age of AIJoin senior risk and payments leaders in Seattle to explore how AI is reshaping fraud strategy. Seats are limited.StripeThe future of agentic commerce is hereExplore how AI agents are transforming commerce at Stripe’s Agentic Commerce Next roadshow. Reserve your spot in Seattle.StripeNew ways to turn global demand into revenueAt Sessions 2026, Stripe unveiled dozens of products and capabilities to help businesses turn global demand into revenue. See how to go global faster with localized checkout and Adaptive Pricing, smarter fraud tools, mul…SalesforceHow Engineering 360 Unified Operations at Scale and Reached 80% AdoptionBy Shiva Nimmagadda, Arun Lakshmi Narayanan, and Arun Gangavarapu. Salesforce engineering teams encountered a significant operational hurdle as the organization scaled. Critical data lived across dozens of fragmented das…DiscordDiscord Patch Notes: June 4, 2026Check out the finer details of the more technical fixes implemented into Discord recently.AirbnbSitar-agent: Building a reliable dynamic configuration sidecar at scaleHow Airbnb built a Kubernetes sidecar to deliver dynamic configuration reliably at scale. By : Bo Teng , Cosmo Qiu , Siyuan Zhou , Ankur Soni , Xin Huang , Willis Harvey Introduction In our previous post , we explored Ai…StripeHelping businesses optimize network costs with the Visa Digital Commerce Authentication Program (DCAP)We moved quickly to help Stripe businesses take advantage of DCAP and capture interchange savings while protecting authorization rates. Here’s what we did.SpotifyCoding Is No Longer the Constraint: Scaling Developer Experience to Teams and Agents at SpotifyAt Code with Claude, Spotify’s chief architect shared how we make both teams and AI agents more effective. The post Coding Is No Longer the Constraint: Scaling Developer Experience to Teams and Agents at Spotify appeared…SalesforceHow Agentforce Conversation Client Accelerated Accessibility Remediation by 5x Using AI-Driven WorkflowsBy Prasanna Krishna Sanagala, Ronak Shah, Sandeep Tailor, and Mani Manjari Velnati. In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight P…NetflixDynamic Repartitioning for Time Series WorkloadsBy Rajiv Shringi , Kaidan Fullerton , Oleksii Tkachuk and Kartik Sathyanarayanan Introduction Netflix’s TimeSeries Abstraction is a scalable system for ingesting and querying petabytes of temporal event data with millise…NetflixDynamic Repartitioning for Time Series WorkloadsBy Rajiv Shringi , Kaidan Fullerton , Oleksii Tkachuk and Kartik Sathyanarayanan Introduction Netflix’s TimeSeries Abstraction is a scalable system for ingesting and querying petabytes of temporal event data with millise…MetaLights Out, Systems On: Validating Instant Power Loss ReadinessWe’re introducing Instantaneous PowerLoss Storm, a new testing paradigm within Meta’s infrastructure for handling and mitigating instant or zero-notice power loss in our data centers. We’re sharing: how we built readines…InstacartSemantic IDs: Product Understanding at ScaleKey Contributors: Shrikar Archak, Karuna Ahuja, Soroush Sobhkhiz, Marko Avdalovic, Xiyu Wang, JiChao Zhang, Hao Yan, Chris Hartley Introduction Operating a grocery catalog at Instacart’s scale means managing millions of…InstacartFrom Scoring to Spelling: Rebuilding Ads Retrieval at InstacartKey Contributors: Karuna Ahuja, Marko Avdalovic, Soroush Sobhkhiz, Shrikar Archak, Xiyu Wang, Ji Chao Zhang, Hao Yan Introduction Every time a user opens Instacart, they see product recommendations: on the retailer home…InstacartSemantic IDs: Product Understanding at ScaleKey Contributors: Shrikar Archak, Karuna Ahuja, Soroush Sobhkhiz, Marko Avdalovic, Xiyu Wang, JiChao Zhang, Hao Yan, Chris Hartley Introduction Operating a grocery catalog at Instacart’s scale means managing millions of…InstacartFrom Scoring to Spelling: Rebuilding Ads Retrieval at InstacartKey Contributors: Karuna Ahuja, Marko Avdalovic, Soroush Sobhkhiz, Shrikar Archak, Xiyu Wang, Ji Chao Zhang, Hao Yan Introduction Every time a user opens Instacart, they see product recommendations: on the retailer home…AirbnbWhen history fails you, borrow from geographyHow Airbnb used sequential geographic recovery signals and prior propagation to generate reliable corridor-level forecasts when local data was scarce. By: Harrison Katz The problem with unprecedented shocks Almost every…NetflixHigh-Throughput Graph Abstraction at Netflix: Part IBy Oleksii Tkachuk , Kartik Sathyanarayanan , Rajiv Shringi Introduction Netflix has a diverse range of graph use cases, each serving specific business needs with unique functionality and performance requirements. These…NetflixHigh-Throughput Graph Abstraction at Netflix: Part IBy Oleksii Tkachuk , Kartik Sathyanarayanan , Rajiv Shringi Introduction Netflix has a diverse range of graph use cases, each serving specific business needs with unique functionality and performance requirements. These…GrabFrom decentralized Docs-as-Code to a centralized repository: Evolving Grab's documentation strategyIntroduction: The journey of documentation at Grab In early 2021, Grab adopted a Docs-as-Code approach to address gaps in our technical documentation processes, as illustrated in our blog post Embracing a Docs-as-Code .…StripeSolo founding is at an all-time high: Top performers have these traits in commonIn 2025, solo founders in the top decile generated 61 times the revenue of the median solo founder in their first six months. We analyzed the data to understand what drives that gap.SlackSlack AI: The Path to Multi-CloudIn early 2023, Slack faced a foundational challenge: serving Large Language Models (LLMs) at enterprise scale with the security, reliability, and performance our customers expect. Over three years, we evolved from basic…Microsoft ResearchData Formulator 0.7: AI-powered data analytics for enterprise dataData Formulator introduces AI-powered analytics for enterprise data workflows. Data teams can easily bring enterprise data into an AI-ready workspace where users can explore, analyze, and visualize data with AI agents to…Mistral AIAI Now Summit 2026Mistral AIVibe gets to work.Mistral AIIntroducing Search ToolkitMicrosoftImprove your agentic developer tools by grounding in Microsoft LearnDevelopment workflows span terminals, IDEs, background agents, and custom assistants. What matters is whether they draw from the same current source. Learn MCP Server gives any MCP-compatible agent direct access to curre…DropboxBeyond code generation: rethinking engineering productivity in the age of AI agentsHow Dropbox is moving from AI tools that assist engineers to agentic systems that can execute scoped tasks, and how we’re building platforms to support those workflows.DiscordOfficial Discord Integrations for Steal a Brainrot, Grow a Garden, Brookhaven RP, and moreHow some of the most popular Roblox games have integrated Discord account linking to enhance social features and safety capabilities for official community servers.YelpBeyond the Menu Tree: How Yelp Built a Smarter Customer Success Chatbot with AIThe Evolution of Support: From Fixed Phrases to Conversation At Yelp, delivering responsive and accurate customer support is a core priority. For years, our legacy Customer Success (CS) Chatbot provided support by guidin…StripeExpanding Stripe Radar to protect more of your businessRadar now blocks high-risk transactions across all supported payment methods; defends against new fraud types like multi-account abuse and pay-as-you-go abuse, regardless of which payment processor you use; and gives pla…SalesforceAgentforce’s Agent Script: Building Deterministic Control for Enterprise AI WorkflowsIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Elijah Ben Izzy, Software Engineering Architect at Salesforce. Elijah is building Age…Microsoft ResearchExtending Human Intelligence Through AIUnderstanding AI as an extension of human intelligence—not a replacement for it—offers a more grounded path for building trustworthy AI systems. The post Extending Human Intelligence Through AI appeared first on Microsof…Mistral AIPhysics AI research that’s shaping the industry.Mistral AIIntroducing physics AI at Mistral: the foundation for engineering acceleration.MicrosoftHow AI coding agents actually use your technologyYou ship an SDK, a CLI, an API, and developers use it. Now AI coding agents use it too, except they use it differently than humans do. Most of the time you have no idea what s actually happening between developer types a…SalesforceBuilding an Enterprise Agent Platform: Enforcing Identity, Data, and API GovernanceWhile enterprises deploy AI agents at a rapid pace, their governance strategies often remain fragmented. Most organizations enforce identity, data access, and API security in separate silos, which creates dangerous gaps…MetaSilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation SystemsWe’re introducing SilverTorch, a reimagining of recommendation systems that unifies all retrieval components for user generated content under a unified architecture. SilverTorch shows up to 23.7x higher throughput compar…EtsyShaping Product Understanding with Contrastive Reinforcement LearningEtsy’s marketplace is defined by the creativity and craftsmanship of our sellers and the hundreds of millions of highly diverse products they offer. You can find silversmiths who cold-forge recycled sterling silver, weav…Mistral AIEmmi joins Mistral to accelerate the AI-native industrySalesforceAgent Fabric Context Catalog and the Future of AI GovernanceModern agents no longer execute within predictable application boundaries. They invoke APIs dynamically, retrieve enterprise context through MCP servers, orchestrate workflows across multiple platforms and interact with…Mistral AIConnect the dots: Build with built-in and custom MCPs in StudioMistral AIRemote agents in Vibe. Powered by Mistral Medium 3.5.Mistral AIConnect the dots: Build with built-in and custom MCPs in StudioMistral AIRemote agents in Vibe. Powered by Mistral Medium 3.5.InstacartHow AI Changes the Role of Applied ScientistsLevi Boxell, Tilman Drerup, Alexandr Lenk The Economics Team at Instacart is an applied science team that operates at the intersection of machine learning engineering and economics. Similar to other applied science teams…InstacartHow AI Changes the Role of Applied ScientistsLevi Boxell, Tilman Drerup, Alexandr Lenk The Economics Team at Instacart is an applied science team that operates at the intersection of machine learning engineering and economics. Similar to other applied science teams…GrabThe Hugo evolution: Engineering Grab's unified, one-click data ingestion platform with Apache FlinkIntroduction Data drives every decision we make at Grab. As our operations scale, so does our need for robust, real-time data ingestion and processing frameworks. Enter Hugo: our self-service data platform that has long…YelpHow Partition Access Visualizations Reduced our Data Lake S3 Cost by 33%Introduction In large analytics environments, data teams often struggle to answer deceptively simple questions, like who their stakeholders are and how their data is being used. At Yelp, we address this by visualizing ac…PinterestMaking User-Sequence Data More Cost-Efficient, Faster, and Easier to UseAuthors ( listed alphabetically ) Ads Feature Engineering Infra team: Ajay Venkatakrishnan, Le Zhang Core ML Infra team: Eric Shang, Pihui Wei ML Data team: Connor Votroubek, Yi He User Understanding team: Camilo Munoz,…Microsoft ResearchVega: Zero-knowledge proofs for digital identity in the age of AIVega turns a full credential into a single proof, sharing only what is needed and nothing more, with performance that works in real apps. The post Vega: Zero-knowledge proofs for digital identity in the age of AI appeare…Microsoft ResearchMagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small modelsMagenticLite is an agentic system for small models that works across the browser and local file system in a single workflow. It combines specialized models and orchestration to support efficient agentic performance on ev…MozillaAnnouncing Web Serial Support in FirefoxSupport for Web Serial in Firefox 151 for Desktop Firefox can now connect directly to microcontrollers, development boards, 3D printers, power meters, and other serial-connected hardware from the web. Starting in Firefox…MicrosoftThe AX stack: what&#8217;s fixed, where you can winAI coding agents promise to make you more productive. On the surface they do, but in practice they fall short: agents generate code that doesn t compile, use a deprecated SDK, or pick the wrong service entirely. Is it yo…DropboxIntroducing Nova, our internal platform for coding agentsNova lets engineers run multiple coding sessions in parallel and lets internal systems use AI agents as part of automated workflows.DiscordMaking It Easier Than Ever to Connect with Friends in League & VAL!Link your Discord and Riot accounts to sync your friends lists, show your in-game activity as your Discord status, and invite your Discord friends directly to your League or VAL lobby.Google DeepMindWe’re launching the Google DeepMind Accelerator program in Asia Pacific to tackle environmental risksYelpOptimizing Our Build Times by Migrating from Webpack to RspackOver the years, Webpack has remained the bundler of choice for many JS projects, including here at Yelp. While it has served us well, its speed has increasingly become a bottleneck as our monorepo continues to grow. Fort…SalesforceHow Salesforce Built an AI Security Agent for Autonomous Threat TriageIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Mor Levi, Vice President of Detection, Analysis and Response at Salesforce, who leads…MicrosoftAgentic-Agile: Why Agent Development Needs Agile (Not Just Prompts)A bad system will beat a good person [or agent] every time ~Dr. William Edwards Deming (with apologies) I started vibe coding by writing prompts (often dictated into my phone), refining them with an agent in M365 Copilot…AirbnbScaling Airbnb’s identity graph with a unified knowledge graph infrastructureHow Airbnb shifts from PaaS to an internal knowledge graph infrastructure at scale. By: Lucen Zhao , Shukun Yang , Ashish Jain Knowledge graphs offer a natural and powerful way to represent relationships between entities…SpotifyBetter Experiments with LLM Evals — A funnel, not a forkTL;DR LLM evals, automated judges that assess relevance, coherence, and quality at scale, are a powerful new... The post Better Experiments with LLM Evals — A funnel, not a fork appeared first on Spotify Engineering .DiscordEvery Voice and Video Call on Discord Is Now End-to-End EncryptedAs of March 2026, E2EE is now enforced for every voice and video call on Discord. This represents a multi-year commitment, and Discord’s VP of Engineering is here to talk about why it matters.Google DeepMindFast-tracking genetic leads to reverse cellular agingBiologists use Co-Scientist to find novel factors that successfully rejuvenate human cells.Google DeepMindMaking it easier to understand how content was created and editedWe're expanding our tools to help you understand how content was created and edited across the web.Google DeepMindGemini for Science: AI experiments and tools for a new era of discoveryA collection of science tools and experiments to expand the scale and precision of scientific exploration.Google DeepMindIntroducing Google Antigravity 2.0Google DeepMindIntroducing Gemini OmniGoogle DeepMindSimulate real-world places with Project Genie and Street ViewWe’re expanding access to Google AI Ultra subscribers globally and introducing a new capability powered by Street View.Google DeepMindAccelerating discovery of liver disease mechanismsFilippo Menolascina uses Co-Scientist to identify new liver disease treatments and explain why existing drugs only help certain patients.Google DeepMindOpening new paths in aging researchCalico Life Sciences uses Co-Scientist to connect scattered findings and generate new leads in aging research.Google DeepMindFinding the molecular switches behind new infectious diseasesClare Bryant uses Co-Scientist to identify genetic triggers in emerging infectious diseases.Google DeepMindStrengthening Singapore’s AI Future: A New National PartnershipGoogle DeepMind and Singapore partner to apply frontier AI to address complex challenges across health, education, and sustainability and more.Microsoft ResearchFurther Notes on Our Recent Research on AI Delegation and Long-Horizon ReliabilityOur recent paper, “LLMs Corrupt Your Documents When You Delegate”, has generated discussion about the reliability of AI systems in delegated workflows. We appreciate the interest in this work and want to clarify several…GrabHow AI is transforming analytics at GrabIntroduction At Grab, analytics sits close to almost every decision that matters. Our north star is the democratization of intelligence, ensuring that anyone making a business call has immediate access to trustworthy ans…GrabScaling developer experience: How we improved Android Studio in a large monorepoIntroduction Long integrated development environment (IDE) sync/indexing times can quietly erode developer productivity, making code navigation sluggish, spiking memory usage, and slowing down Jetpack Compose preview upd…SalesforceCreating a Multi-Tenant AI Agent Platform Handling 7K+ Sessions Without Cross-Team InterferenceIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Priyanka Saraf, Senior Software Engineer on the Agentforce Foundations team. Priyanka…InstacartScaling Personalized Marketing for Multi-Tenant Commerce PlatformsTL;DR Background: Marketing Across Marketplace and Storefront Instacart operates across two distinct commerce experiences: Instacart Marketplace, our first-party consumer marketplace Storefront Pro, our white-label e-com…InstacartScaling Personalized Marketing for Multi-Tenant Commerce PlatformsTL;DR Background: Marketing Across Marketplace and Storefront Instacart operates across two distinct commerce experiences: Instacart Marketplace, our first-party consumer marketplace Storefront Pro, our white-label e-com…Microsoft Researchmimalloc: A new, high-performance, scalable memory allocator for the modern eramimalloc is an open-source, modern, scalable memory allocator that is a drop-in replacement for malloc and free. It is relatively small (~12K lines), with clear internal data structures, and is easy to build and integrat…Microsoft ResearchGridSFM: A new, small foundation model for the electric gridIntroducing GridSFM, a small foundation model that can predict AC optimal power flow in milliseconds, boosting efficiency and unlocking cost savings. Learn how GridSFM gives grid operators direct visibility into congesti…MetaReel Friends: Building Social Discovery that Scales to BillionsOn its face the new Friend Bubbles feature looks simple enough. It highlights Reels your friends have watched and reacted to. But sometimes the features that seem the most straightforward require the deepest engineering…DiscordCelebrate Discord’s 11th Birthday with an Exclusive Set of Emoji and WallpapersDiscord is turning 11 this year! To celebrate, we made over twenty Discord-themed emojis, along with over twenty wallpapers, and even a digital poster for everyone to download, for free!AirbnbViaduct 1.0 and the future of Airbnb’s data meshMoving from an internal tool to a community-driven, production-ready data mesh. By : Ryan Tanner , Raymie Stata , Adam Miskiewicz Introduction We’re excited to announce the 1.0 release of the Viaduct. This release marks…PinterestAn Engineer’s Guide to Better AI Skills: Implementing a Testing Process to Optimize Agent…An Engineer’s Guide to Better AI Skills: Implementing a Testing Process to Optimize Agent Performance in Any Repository or Skill Author: Daniel Reed The tech industry is currently seeing a massive overhaul in the way we…Microsoft ResearchAdvancing AI for materials with MatterSim: experimental synthesis, faster simulation, and multi-task modelsMatterSim is expanding what AI can do for materials science—from faster large-scale simulations to MatterSim-MT, a new multi-task model for simulating properties beyond potential energy surfaces alone. The post Advancing…MetaMigrating Data Ingestion Systems at Meta ScaleMeta’s data ingestion system, which our engineering teams leverage for up-to-date snapshots of the social graph, has recently undergone a significant revamp to enhance its reliability at scale. Moving from our legacy sys…YelpML based ranking using NrtsearchWe’ve extended Nrtsearch with the Inference Plugin, which embeds ML-based ranking directly in the search layer — eliminating the need for a standalone scoring service. We use Nrtsearch (read more information on the blog…StripeFive vertical SaaS insights from Sessions 2026AI is forcing platforms to expand beyond pure software. See how vertical SaaS platforms are using payments, financial services, and agentic commerce to build more durable businesses.Microsoft ResearchSocialReasoning-Bench: Measuring whether AI agents act in users’ best interestsUsing SocialReasoning Bench, we observed a stable pattern across models—agents execute competently, but fail to consistently improve the user’s position, even with explicit instructions to optimize for user interest. The…MetaLabyrinth 1.1: Making End-to-End Encrypted Backups Even More ReliableWe’re rolling out version 1.1 of Labyrinth, the encrypted storage system and protocol that secures messages and history on Messenger. Labyrinth 1.1 enhances the reliability of end-to-end encrypted backups with a new sub-…DiscordHow to Use Nitro: A Beginner’s Guide to Discord’s Premium SubscriptionWhat’s Discord Nitro all about? What perks does it give, and how can you get it? If you’re looking to expand your Nitro knowledge, you’re in the right place.DiscordNitro Now Comes with Xbox Game Pass and New Benefits. Welcome to Nitro Rewards.As we hit Nitro’s 10-year anniversary, we re launching Nitro Rewards: a brand-new benefits program built with some of the biggest names in gaming. See what’s coming for Nitro members, for no added cost.PinterestEnhancing Ad Relevance: Integrating Real-Time Context into Sequential Recommender ModelsHuiqin Xin | Machine Learning Engineer II, Ads Vertical Modeling; Lakshmi Manoharan | Senior Machine Learning Engineer, Ads Vertical Modeling; Karthik Jayasurya | Staff Machine Learning Engineer, Ads Signals; Ziwei Guo |…NetflixScaling ArchUnit with Nebula ArchRulesBy John Burns and Emily Yuan Introduction At Netflix, we operate using a polyrepo strategy with tens of thousands of Java repositories. This means that we need to have ways of sharing common build logic across these repo…NetflixScaling ArchUnit with Nebula ArchRulesBy John Burns and Emily Yuan Introduction At Netflix, we operate using a polyrepo strategy with tens of thousands of Java repositories. This means that we need to have ways of sharing common build logic across these repo…DiscordHow Discord Automates ScyllaDB Clusters at ScaleYou ve been asked to stand up a brand-new database cluster, meaning a whole day of configuring dozens of nodes, validating replication, wiring up dual-write pipelines… what if this whole ordeal took less than two hours?…BAIR (Berkeley)Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling.apr-fig { text-align: center; margin: 1.35em 0; line-height: 1.4; } .apr-fig--wide img { display: inline-block; width: 100%; max-width: 100%; height: auto; vertical-align: middle; } .apr-fig--wide-0-8 { max-width: 80%;…MozillaBehind the Scenes Hardening Firefox with Claude Mythos PreviewTwo weeks ago we announced that we had identified and fixed an unprecedented number of latent security bugs in Firefox with the help of Claude Mythos Preview and other AI models. In this post, we’ll go into more detail a…GrabEnhancing Flink deployment with shadow testingIntroduction Ensuring the reliability of Apache Flink deployments in Grab is crucial for the availability of our business-critical, real-time applications. While all applications are tested in a staging environment befor…DiscordStock Up in the New Rust Shop! Enjoy a Discord-Only 20% Sale on Most Items until 5/21Starting today, you can now browse, purchase, and even gift and wishlist in-game items for Rust directly in Discord! See how it all works, and learn about a hefty two-week Discord-exclusive launch discount on a wide sele…SlackFrom SSH to REST: A Security-Driven Modernization of Slack&#8217;s EMR Data PipelinesExcerpt By 2024, Slack s data platform had accumulated 700+ SSH-based operators orchestrating critical data pipelines. We re talking daily search indexing that processed terabytes of data, analytics jobs powering busines…MozillaTrustworthy JavaScript for the Open WebThe open web is a critical platform for applications that handle highly sensitive data, from private communications to financial transactions and medical records. Traditionally, servers are trusted to deliver the appropr…MicrosoftAzure Cosmos DB Conf 2026 Recap: Lessons from ProductionA team was running at 100% RU utilization. Throttles were compounding into retries. P99 latency was degrading. The assumption was obvious: provision more throughput. They didn’t. Instead, they found a single logical part…AirbnbMonitoring reliably at scaleDesigning monitoring that works when everything else doesn’t. By : Abdurrahman J. Allawala Introduction When an incident hits, teams lean on observability to answer the only questions that matter: what’s broken, and why?…NetflixDemocratizing Machine Learning at Netflix: Building the Model Lifecycle GraphSaish Sali , Nipun Kumar , Sura Elamurugu Introduction As Netflix has grown, machine learning continues to support our ability to deliver value to members and drive excellence across multiple areas of our business. When…NetflixDemocratizing Machine Learning at Netflix: Building the Model Lifecycle GraphSaish Sali , Nipun Kumar , Sura Elamurugu Introduction As Netflix has grown, machine learning continues to support our ability to deliver value to members and drive excellence across multiple areas of our business. When…InstacartEmpowering Carrot Ads with Domain Adaptive LearningAuthors: Trey Zhong, Xiyu Wang Contributors: Joseph Haraldson, Sharad Gupta, Sarah Lamacchia Introduction Carrot Ads is Instacart’s omnichannel retail media solution that allows retailer partners to build and scale their…InstacartEmpowering Carrot Ads with Domain Adaptive LearningAuthors: Trey Zhong, Xiyu Wang Contributors: Joseph Haraldson, Sharad Gupta, Sarah Lamacchia Introduction Carrot Ads is Instacart’s omnichannel retail media solution that allows retailer partners to build and scale their…DiscordDiscord Patch Notes: May 4, 2026Check out the finer details of the more technical fixes implemented into Discord recently.SpotifyBuilding a Natural Language Interface to the Spotify Ads API with Claude Code PluginsTurning OpenAPI spec and Markdown files into a conversational ads management tool — no compiled code required. The post Building a Natural Language Interface to the Spotify Ads API with Claude Code Plugins appeared first…PinterestOptimizing ML Workload Network Efficiency (Part I): Feature TrimmerGuangtong Bai | Staff Software Engineer, Product ML Infrastructure*; Shantam Shorewala | Software Engineer II, Product ML Infrastructure*; Chi Zhang | Staff Software Engineer, AI Platform*; Neha Upadhyay | Software Engin…NetflixState of Routing in Model ServingBy Nipun Kumar , Rajat Shah , Peter Chng Introduction This is the first blog post in a multi-part series that shares technical insights into how our ML model serving infrastructure powers several personalized experiences…NetflixState of Routing in Model ServingBy Nipun Kumar , Rajat Shah , Peter Chng Introduction This is the first blog post in a multi-part series that shares technical insights into how our ML model serving infrastructure powers several personalized experiences…MetaHow Meta Is Strengthening End-to-End Encrypted BackupsThe HSM-based Backup Key Vault Meta s HSM-based Backup Key Vault provides the foundation for end-to-end encrypted backups for WhatsApp and Messenger. The system allows people to protect their backed-up message history wi…GrabData Mesh at Grab (Part II): The foundational tools behind certificationIntroduction In Part I , we discussed why Grab is investing in a data mesh, referred to as the Signals Marketplace within Grab, as part of our evolving data culture. We also explained how data certification aids teams in…StripeEverything we announced at Sessions 2026We’re making Stripe even more programmable; protecting and propelling your business with the strength of the Stripe network; and building economic infrastructure for AI.StripeGiving agents the ability to payLink’s wallet for agents gives agents programmatic access to Link, including the ability to generate a one-time-use card or Shared Payment Token (SPT) backed by the cards and bank accounts already in your wallet. It’s bu…DiscordYou’ve Got (Too Much) Mail: Behind the Scenes of the 3/25/26 Voice OutageOn March 25th, voice and video on Discord suffered major degradation beginning at 12:13 PDT, lasting a little over three hours. Learn how the issue originated, how it affected systems across Discord, how we recovered, an…AirbnbSkipper: Building Airbnb’s embedded workflow engineHow Airbnb built a lightweight workflow engine to solve durable execution. By : Ricardo Gamba , Andriy Sergiyenko Introduction: The durable execution problem Picture this hypothetical flow: A host submits an insurance cl…PinterestFrom Clicks to Conversions: Architecting Shopping Conversion Candidate Generation at PinterestAuthors: Richard Huang | Machine Learning Engineer II; Yu Liu | Senior Machine Learning Engineer; Ziwei Guo | Senior Machine Learning Engineer; Andy Mao | Staff Machine Learning Engineer; Supeng Ge | Sr. Staff Machine Le…Mistral AIWorkflows for work that runs the businessNetflixScaling Camera File Processing at NetflixOrchestrating Media Workflows Through Strategic Collaboration Authors: Eric Reinecke , Bhanu Srikanth Introduction to Content Hub’s Media Production Suite At Netflix, we want to provide filmmakers with the tools they nee…NetflixScaling Camera File Processing at NetflixOrchestrating Media Workflows Through Strategic Collaboration Authors: Eric Reinecke , Bhanu Srikanth Introduction to Content Hub’s Media Production Suite At Netflix, we want to provide filmmakers with the tools they nee…DiscordMeasure Less to Learn More: Using Fewer, Higher-quality Metrics to Capture What MattersToo many experiment metrics can make meaningful changes harder to detect. Learn how Discord used simulations and Principal Component Analysis to maximize signal and reduce noise.MicrosoftLangChain.js for Beginners: A Free Course to Build Agentic AI Apps with JavaScriptWant to build AI agents with JavaScript that go beyond basic chat completions? Agents that reason, call tools, and pull from knowledge bases on their own? We put together a free, open source course to help you get there.…LyftHow We Built a Smarter Pickup Experience for Gated CommunitiesIf you live in a gated community, you’ve been there: You request a ride from your apartment complex, expect your driver to come to you as usual, and then — your driver’s car icon just stops right at the front gate. You w…LyftHow We Built a Smarter Pickup Experience for Gated CommunitiesIf you live in a gated community, you’ve been there: You request a ride from your apartment complex, expect your driver to come to you as usual, and then — your driver’s car icon just stops right at the front gate. You w…YelpHow Yelp Keeps Server-Driven UI Consistent Across Four PlatformsIf you’ve read our earlier post, you already know about CHAOS—the server-driven UI (SDUI) framework we built at Yelp that powers our dynamic views. Until now, we’ve explored its architecture, backend implementation, and…SpotifyBackground Coding Agents: Supercharging Downstream Consumer Dataset Migrations (Honk, Part 4)How we used Honk, Backstage, and Fleet Management to ease the pain of migrating thousands of datasets. The post Background Coding Agents: Supercharging Downstream Consumer Dataset Migrations (Honk, Part 4) appeared first…MetaModernizing the Facebook Groups Search to Unlock the Power of Community KnowledgeWe’ve fundamentally transformed Facebook Groups Search to help people more reliably discover, sort through, and validate community content that’s most relevant to them. We’ve adopted a new hybrid retrieval architecture a…AirbnbBuilding a fault-tolerant metrics storage system at AirbnbHow we built a storage system that ingests 50 million samples per second and stores 2.5 petabytes of logical time series data. By : Rishabh Kumar Modern observability practice encourages instrumenting every meaningful co…PinterestSmarter URL Normalization at Scale: How MIQPS Powers Content Deduplication at PinterestShanhai Liao | Senior Software Engineer, Content Acquisition and Media Platform; Di Ruan, | Senior Staff Software Engineer, Content Acquisition and Media Platform; Evan Li, | Senior Engineering Manager, Content Acquisiti…BAIR (Berkeley)Gradient-based Planning for World Models at Longer Horizons.grasp-results-table table { font-size: 0.875rem; line-height: 1.35; width: 100%; } .grasp-results-table th, .grasp-results-table td { padding: 0.35rem 0.5rem; } /* Consistent whitespace between major sections (this post…NetflixThe Human Infrastructure: How Netflix Built the Operations Layer Behind Live at ScaleBy: Brett Axler , Casper Choffat , and Alo Lowry In the three years since our first Live show, Chris Rock: Selective Outrage , we have witnessed an incredible expansion of our live content slate and the live operations t…NetflixThe Human Infrastructure: How Netflix Built the Operations Layer Behind Live at ScaleBy: Brett Axler , Casper Choffat , and Alo Lowry In the three years since our first Live show, Chris Rock: Selective Outrage , we have witnessed an incredible expansion of our live content slate and the live operations t…MetaPost-Quantum Cryptography Migration at Meta: Framework, Lessons, and TakeawaysWe’re sharing lessons learned from Meta’s post-quantum cryptography (PQC) migration to help other organizations strengthen their resilience as industry transitions to post-quantum cryptography standards. We’re proposing…MetaCapacity Efficiency at Meta: How Unified AI Agents Optimize Performance at HyperscaleWe re sharing insights into Meta s Capacity Efficiency Program, where we ve built an AI agent platform that helps automate finding and fixing performance issues throughout our infrastructure. By leveraging encoded domain…DiscordMaking Discord on Desktop Look Just Right: Display Settings to Ease the EyesLearn all sorts of toggles, options, and features on Discord’s desktop app to help you view media at your pace, lower the strength of colors across the app, and make app content easier to see.PinterestFinding zombies in our systems: A real-world story of CPU bottlenecksVaibhav Shankar; Staff Software Engineer | Raymond Lee; Staff Software Engineer | Chia-Wei Chen; Staff Software Engineer | Shunyao Li; Sr. Software Engineer | Yi Li; Staff Software Engineer | Ambud Sharma; Principal Engi…AirbnbPrivacy-first connections: Empowering social experiences at AirbnbDiscover how Airbnb prioritizes user privacy while building a more connected community, empowering guests to engage socially, connect confidently, and maintain control of their personal data. By: Joy Jing ✨ Building a mo…SlackManaging context in long-run agentic applicationsExcerpt In complex, long-running agentic systems, maintaining alignment and coherent reasoning between agents requires careful design. In this second article of our series, we explore these challenges and the mechanisms…PinterestScaling Recommendation Systems with Request-Level DeduplicationAuthors: Matt Lawhon | Sr. Machine Learning Engineer; Filip Ryzner | Machine Learning Engineer II; Kousik Rajesh | Machine Learning Engineer II; Chen Yang | Sr. Staff Machine Learning Engineer; Saurabh Vishwas Joshi | Pr…NetflixEvaluating Netflix Show Synopses with LLM-as-a-Judgeby Gabriela Alessio , Cameron Taylor , and Cameron R. Wolfe Introduction When members log into Netflix, one of the hardest choices is what to watch. The challenge isn’t a lack of options — there are thousands of titles —…NetflixEvaluating Netflix Show Synopses with LLM-as-a-Judgeby Gabriela Alessio , Cameron Taylor , and Cameron R. Wolfe Introduction When members log into Netflix, one of the hardest choices is what to watch. The challenge isn’t a lack of options — there are thousands of titles —…PinterestPerformance for EveryoneAuthor: Lin Wang (Android Performance Engineer) Default Feature For mobile apps, performance is considered as the “default feature”, which means apps are expected to run fast and be responsive. It’s just as if we expect…YelpZero downtime Upgrade: Yelp’s Cassandra 4.x Upgrade StoryThe Database Reliability Engineering team at Yelp seamlessly upgraded more than a thousand Cassandra nodes with zero downtime. This post takes you behind the scenes of our upgrade strategy, from planning sessions to flaw…PinterestEvolution of Multi-Objective Optimization at Pinterest Home feedHomefeed: Jiacong He, Dafang He, Jie Cheng (former), Andreanne Lemay, Mostafa Keikha, Rahul Goutam, Dhruvil Deven Badani, Dylan Wang Content Quality: Jianing Sun, Qinglong Zeng ML Serving: Li Tang Introduction In feed re…AirbnbBuilding a high-volume metrics pipeline with OpenTelemetry and vmagentA production-tested approach for moving a large-scale metrics pipeline from StatsD to OpenTelemetry and Prometheus. By: Eugene Ma , Natasha Aleksandrova When migrating to a new monitoring system, you’ll want to frontload…NetflixStop Answering the Same Question Twice: Interval-Aware Caching for Druid at Netflix ScaleBy Ben Sykes In a previous post , we described how Netflix uses Apache Druid to ingest millions of events per second and query trillions of rows, providing the real-time insights needed to ensure a high-quality experienc…NetflixStop Answering the Same Question Twice: Interval-Aware Caching for Druid at Netflix ScaleBy Ben Sykes In a previous post , we described how Netflix uses Apache Druid to ingest millions of events per second and query trillions of rows, providing the real-time insights needed to ensure a high-quality experienc…DiscordDiscord Patch Notes: April 6, 2026Check out the finer details of the more technical fixes implemented into Discord recently.DropboxImproving storage efficiency in Magic Pocket, our immutable blob storeBy turning compaction into a layered, adaptive pipeline and strengthening our monitoring and controls, we made Magic Pocket more resilient to workload changes.DiscordMULTIPLAYER SEQUEL TO ACCLAIMED AAAA GAME “THE LAST MEADOW” ANNOUNCED: PLAYABLE NOWBand together with Discordians from across the world in Last Meadow Online, the world’s first DBMMIRPG. Available to play until April 7, 2026.SlackFrom Custom to Open: Scalable Network Probing and HTTP/3 Readiness with PrometheusThe Problem: Legacy Tooling and Its Limitations Currently, Slack utilizes a hybrid approach to network measurement, incorporating both internal (such as traffic between AWS Availability Zones) and external (monitoring tr…Mistral AISpaces: A CLI Built for Humans and AgentsLyftPredicting Rider Conversion in Sparse Data Environments with Bayesian TreesAt Lyft, understanding how riders go through our user experience is fundamental to operating a healthy marketplace. Specifically, it is important to have a robust model determining if a rider will actually request a ride…LyftPredicting Rider Conversion in Sparse Data Environments with Bayesian TreesAt Lyft, understanding how riders go through our user experience is fundamental to operating a healthy marketplace. Specifically, it is important to have a robust model determining if a rider will actually request a ride…YelpBuilding Biz Ask Anything: From Prototype to ProductIntroduction Users have access to a wealth of information on Yelp business pages – from reviews and photos to structured information, menus, and Ask the Community feature on the business page, a single business page can…DiscordHow Multi-Factor Authentication Helps Keep Your Discord Account SafeA Discord account is more than just your username and avatar. That’s why it’s important to help keep your account safe and secure by using Multi-Factor Authentication, SMS Backup Authentication & QR Code Login. Learn how…MozillaFirefox Developer Edition and Beta: Try out Mozilla&#8217;s .rpm package!In January, we introduced our Nightly package for RPM-based Linux distributions. Today, we are thrilled to announce it is now available for Firefox Beta! Firefox Beta is great for testing your sites in a version of Firef…LyftBeyond A/B Testing: Using Surrogacy and Region-Splits to Measure Long-Term Effects in MarketplacesImage generated with Gemini 3 Pro (Google), 2026. Written by Amber Wang and Y oonji Kim at Lyft. Background Whenever you use the Lyft app, there is a complex balancing act happening behind the scenes. Various levers are…LyftBeyond A/B Testing: Using Surrogacy and Region-Splits to Measure Long-Term Effects in MarketplacesImage generated with Gemini 3 Pro (Google), 2026. Written by Amber Wang and Y oonji Kim at Lyft. Background Whenever you use the Lyft app, there is a complex balancing act happening behind the scenes. Various levers are…DropboxReducing our monorepo size to improve developer velocityMonorepos will continue to grow as products evolve, but growth doesn’t have to mean friction.DiscordDiscord Update: March 24, 2026 ChangelogHere s the Discord Changelog from March 24, 2026, so you can stay informed on what’s new in recent app updates!Mistral AISpeaking of VoxtralEtsyMaking Ads Count: Using MMoE and Auxiliary Tasks to Better Connect Buyers & SellersWhen buyers search on Etsy, they need to quickly and easily find the perfect item. At the same time, sellers need to be confident their unique products are being seen by the right customers. Our Ads Search ranking model,…SlackHow Slack Rebuilt Notifications 📣Introduction 🔔 At Slack, notifications are how teams stay in the loop, but they can also become overwhelming when not designed with intention. Our goal was to make staying informed feel effortless. We set out to rebuild…GrabFrom firefighting to building: How AI agents restored our team’s core productivityAbstract Grab’s Analytics Data Warehouse (ADW) team supports over 1,000 users each month. These users support an extensive repository of more than 15,000 tables, which powers approximately 50% of all queries within our d…EtsyMigrating Etsy’s database sharding to VitessEtsy has maintained a sharded MySQL architecture since around 2010. This database cluster contains most of Etsy’s online data and is made up of ~1,000 tables distributed across ~1,000 shards. Over the last 16 years, it h…Mistral AIIntroducing ForgeDropboxHow we optimized Dash&#039;s relevance judge with DSPyWe used DSPy to turn prompt engineering for our relevance judge into a measurable, automated optimization loop, improving task performance, cost, and how reliably it works in production.Mistral AIMistral AI partners with NVIDIA to accelerate open frontier modelsMistral AILeanstral: Open-Source foundation for trustworthy vibe-codingMistral AIIntroducing Mistral Small 4DiscordHow ROOST is Advancing Online SafetyThe threat landscape online has shifted dramatically. Many online platforms are left to reinvent safety tools from scratch. That’s the gap ROOST was built to close — and it’s why open-sourcing battle-tested tools like Os…BAIR (Berkeley)Identifying Interactions at Scale for LLMs--> Understanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence. Interpretability research aims to make the decisio…GrabEnabling R8 optimization at scale with AI-assisted debuggingGrab is Southeast Asia’s leading superapp, providing a suite of services that bring essential needs to users throughout the region. Its offerings include ride-hailing, food delivery, parcel delivery, mobile payments, and…DiscordYou’re Now Discord Official: Developers, Claim Your Game and Verify Your ServerDevelopers can now claim and customize their game’s profiles on Discord. Curate your game’s presence on the platform to help people discover more about your game, and get your server verified in the process! Read on to s…Mistral AIRails testing on autopilot: Building an agent that writes what developers won'tDiscordBuilding on the Social Layer of Games: What’s New from GDC 2026At GDC 2026, Discord gives developers more ways to close the gap between connection and play.GrabReclaiming Terabytes: Optimizing Android image caching with TLRUIntroduction In a previous post, we discussed Project Bonsai , our initiative to reduce the Grab app’s download size. We successfully reduced the Android Application Package (APK) download size by 26%. This reduction off…DiscordDiscord Patch Notes: March 6, 2026Check out the finer details of the more technical fixes implemented into Discord recently.DiscordTracing Discord&#39;s Elixir Systems (Without Melting Everything)Join Senior Software Engineer Nick Krichevsky as he explains how Discord added distributed tracing to Elixir s message passing and optimized it to handle millions of concurrent users.MozillaWhy is WebAssembly a second-class language on the web?This post is an expanded version of a presentation I gave at the 2025 WebAssembly CG meeting in Munich. WebAssembly has come a long way since its first release in 2017. The first version of WebAssembly was already a grea…InstacartOur Early Journey to Transform Instacart’s Discovery Recommendations with LLMsKey Contributors: Moein Hasani, Hamidreza Shahidi, Trace Levinson, Guanghua Shu Introduction At Instacart, we are laser-focused on improving the user experience by making shopping feel easy, engaging, and personalized. O…InstacartOur Early Journey to Transform Instacart’s Discovery Recommendations with LLMsKey Contributors: Moein Hasani, Hamidreza Shahidi, Trace Levinson, Guanghua Shu Introduction At Instacart, we are laser-focused on improving the user experience by making shopping feel easy, engaging, and personalized. O…DropboxUsing LLMs to amplify human labeling and improve Dash search relevanceHow we train Dash's search ranking models with a mix of human and LLM-assisted labeling.MozillaGoodbye innerHTML, Hello setHTML: Stronger XSS Protection in Firefox 148Cross-site scripting (XSS) remains one of the most prevalent vulnerabilities on the web. The new standardized Sanitizer API provides a straightforward way for web developers to sanitize untrusted HTML before inserting it…DiscordGetting Global Age Assurance Right: What We Got Wrong and What&#39;s ChangingDiscord s CTO addresses community concerns about age assurance: no mass ID collection, new vendor transparency commitments, and a delayed global launch until second half of 2026.LyftScaling Localization with AI at LyftWritten by Stefan Zier For years, Lyft’s localization infrastructure relied exclusively on human translation. While this model usually ensured excellent quality, it was bound by multi-day turnarounds and costs that scale…LyftScaling Localization with AI at LyftWritten by Stefan Zier For years, Lyft’s localization infrastructure relied exclusively on human translation. While this model usually ensured excellent quality, it was bound by multi-day turnarounds and costs that scale…DiscordHow to Change Your Theme to Bring Your Vibe to DiscordAdd a splash of personality and make Discord pop by changing Discord’s color scheme! Learn how to adjust the look of Discord on both desktop and mobile.DiscordOsprey: Open Sourcing our Rule EngineDiscord uses Osprey to quickly detect and remove new types of harm from putting our customers at risk. Now we’re open-sourcing this tool so others can do the same.InstacartTurning Data into Velocity: Caper’s Edge and Cloud Data Flywheel with CapsightKey Contributors: Youming Luo, Andrew Tanner, Matas Sriubiskis, Sylvia Lin, Sikun Zhu, Lei Li, Xiao Zhou Introduction Caper is Instacart’s AI-powered smart cart that provides customers with a fast, seamless, and intuitiv…InstacartTurning Data into Velocity: Caper’s Edge and Cloud Data Flywheel with CapsightKey Contributors: Youming Luo, Andrew Tanner, Matas Sriubiskis, Sylvia Lin, Sikun Zhu, Lei Li, Xiao Zhou Introduction Caper is Instacart’s AI-powered smart cart that provides customers with a fast, seamless, and intuitiv…MozillaLaunching Interop 2026The Interop Project is a cross-browser initiative to improve web compatibility in areas that offer the most benefit to both users and developers. The group, including Apple, Google, Igalia, Microsoft, and Mozilla, takes…LyftTrusting the Untestable: Validation and Diagnostics for the Doubly Robust Modelswritten by Ross Chu and Shima Nassiri The Causal Frontier: Measurement Beyond Randomization The gold standard for determining the causal impact of a policy or product change at a company like Lyft is the A/B test (random…LyftTrusting the Untestable: Validation and Diagnostics for the Doubly Robust Modelswritten by Ross Chu and Shima Nassiri The Causal Frontier: Measurement Beyond Randomization The gold standard for determining the causal impact of a policy or product change at a company like Lyft is the A/B test (random…DropboxHow low-bit inference enables efficient AIMaking products like Dropbox Dash accessible to individuals and businesses means tackling new challenges around efficiency and resource use.DropboxInsights from our executive roundtable on AI and engineering productivityFrom Claude Code to Cursor, we're big adopters of AI coding tools at Dropbox. The early results have been promising, but there are still a lot of open questions about how to work with these tools most effectively and whe…InstacartFrom print to digital: Making weekly flyers shoppable at Instacart through computer vision and LLMsFrom Print to Digital: Making Weekly Flyers Shoppable at Instacart Through Computer Vision and LLMs Key contributors: Prithvi Srinivasan, Shishir Kumar Prasad, Kristen Morgan, Bryan Pham, Rick Shukla, Preeti Chadha, Vipu…InstacartFrom print to digital: Making weekly flyers shoppable at Instacart through computer vision and LLMsFrom Print to Digital: Making Weekly Flyers Shoppable at Instacart Through Computer Vision and LLMs Key contributors: Prithvi Srinivasan, Shishir Kumar Prasad, Kristen Morgan, Bryan Pham, Rick Shukla, Preeti Chadha, Vipu…Mistral AIVoxtral transcribes at the speed of sound.Mistral AIVoxtral transcribes at the speed of sound.DiscordDiscord Patch Notes: February 4, 2026Check out the finer details of the more technical fixes implemented into Discord recently.InstacartMigrating to Jetpack ComposeMigrating to Jetpack Compose: How AI Accelerated Our Journey at Caper Introduction At Instacart, our Caper smart carts bring together AI, computer vision, and real-time data to power the future of in-store shopping. Cust…InstacartMigrating to Jetpack ComposeMigrating to Jetpack Compose: How AI Accelerated Our Journey at Caper Introduction At Instacart, our Caper smart carts bring together AI, computer vision, and real-time data to power the future of in-store shopping. Cust…YelpHow Yelp Built a Back-Testing Engine for Safer, Smarter Ad Budget AllocationIntroduction Modern advertising platforms are fast-paced and interconnected: even small adjustments can have ripple effects on how ads are shown, how budgets are spent, and the value advertisers get from their ad spend.…DiscordHow to Customize Your Discord ProfileMake your first impression count with a profile that represents you how YOU want to be seen. Learn how to edit and customize your profile to have it rep you the right way.GrabCursor at Grab: Adoption and impactAdoption overview The illustration below encapsulates how Cursor is scaled across Grab, achieving rapid and widespread adoption that accelerated software development and empowered non-technical teams to build solutions.…DropboxEngineering VP Josh Clemm on how we use knowledge graphs, MCP, and DSPy in DashEngineering VP Josh Clemm deep-dives into how we think about knowledge graphs, indexes, MCP, and prompt optimization using tools like DSPy.Mistral AITerminally online Mistral Vibe.Mistral AITerminally online Mistral Vibe.Mistral AIHeaps do lie: debugging a memory leak in vLLM.EtsyHow Etsy Uses LLMs to Improve Search RelevanceEver searched for something specific, only to be met with results that are close, but not quite ? On Etsy’s Search Relevance team, that frustration is exactly what we are tackling. Our goal is simple yet ambitious: to he…BAIR (Berkeley)Information-Driven Design of Imaging SystemsAn encoder (optical system) maps objects to noiseless images, which noise corrupts into measurements. Our information estimator uses only these noisy measurements and a noise model to quantify how well measurements disti…LyftLyft’s Feature Store: Architecture, Optimization, and EvolutionWritten by Rohan Varshney , with support from Devon Mittow Janice Lee . This article expands upon a presentation from the Feature Store Summit 2025, which can be viewed in full here . There is also another video availabl…LyftLyft’s Feature Store: Architecture, Optimization, and EvolutionWritten by Rohan Varshney , with support from Devon Mittow Janice Lee . This article expands upon a presentation from the Feature Store Summit 2025, which can be viewed in full here . There is also another video availabl…Mistral AIIntroducing Mistral OCR 3LyftFrom Python3.8 to Python3.10: Our Journey Through a Memory LeakImage generated with ChatGPT (OpenAI), 2025. Intro When working with Python, memory management often feels like a solved problem. The garbage collector quietly does its job, and unlike C or C++, we rarely think about mal…LyftFrom Python3.8 to Python3.10: Our Journey Through a Memory LeakImage generated with ChatGPT (OpenAI), 2025. Intro When working with Python, memory management often feels like a solved problem. The garbage collector quietly does its job, and unlike C or C++, we rarely think about mal…DiscordHow to Make and Use Custom Emoji on DiscordEmojis on Discord are special — you can make a little picture out of almost any symbol, in-joke, or bizarre late-night inspiration.Mistral AIIntroducing: Devstral 2 and Mistral Vibe CLI.Mistral AIIntroducing Mistral 3SlackStreamlining Security Investigations with AgentsSlack’s Security Engineering team is responsible for protecting Slack’s core infrastructure and services. Our security event ingestion pipeline handles billions of events per day from a diverse array of data sources. Rev…EtsyReducing experiment duration with predicted control variatesI n 2021, we published a blog post titled “ Increasing experimentation accuracy and speed by using control variates ,” describing how we reduce the variance of metrics using CUPED in our experimentation platform. This is…SlackAndroid VPAT journeyBackground A Voluntary Product Accessibility Template (VPAT) is a document that outlines how well a product aligns with accessibility (a11y) standards. Its primary purpose is to inform customers about a product s a11y fe…Mistral AIMistral AI - KI für DeutschlandLyftLyftLearn Evolution: Rethinking ML Platform ArchitectureWritten by Yaroslav Yatsiuk At Lyft, machine learning (ML) is the engine behind our most critical business functions — from dispatch and pricing optimization to fraud detection and support automation. Our ML infrastructu…LyftLyftLearn Evolution: Rethinking ML Platform ArchitectureWritten by Yaroslav Yatsiuk At Lyft, machine learning (ML) is the engine behind our most critical business functions — from dispatch and pricing optimization to fraud detection and support automation. Our ML infrastructu…InstacartBuilding The Intent Engine: How Instacart is Revamping Query Understanding with LLMsAuthors: Yuanzheng Zhu, Guanghua Shu, Raochuan Fan, Vinesh Gudla, Tejaswi Tenneti Introduction When people search for items on Instacart, they don’t always type perfectly worded phrases. They might write “bread no gluten…InstacartBuilding The Intent Engine: How Instacart is Revamping Query Understanding with LLMsAuthors: Yuanzheng Zhu, Guanghua Shu, Raochuan Fan, Vinesh Gudla, Tejaswi Tenneti Introduction When people search for items on Instacart, they don’t always type perfectly worded phrases. They might write “bread no gluten…BAIR (Berkeley)RL without TD learningIn this post, I’ll introduce a reinforcement learning (RL) algorithm based on an “alternative” paradigm: divide and conquer . Unlike traditional methods, this algorithm is not based on temporal difference (TD) learning (…EtsyImproving performance by prefetching product pages from Etsy SearchRarely are there opportunities for big, bold, game-changing improvements in web performance. The Speculation Rules API (SRA) is a recent browser development that offers just such an opportunity. This post details a joint…Mistral AIIntroducing Mistral AI Studio.EtsyUnderstanding Etsy’s Vast Inventory with LLMsFor more than 20 years, Etsy has been the destination for human creativity online. Our marketplace is home to more than 100 million special items made, handpicked and designed by more than 5 million sellers. These items…EtsyUnlocking Faster Insights with Experimenter-Defined SegmentationsImagine you have a fabulous idea to drive more sales on Etsy by giving out free ice cream with every purchase. How would you know if it will actually work? One way to test this out is to run an experiment ! An experiment…YelpS3 server access logs at scaleIntroduction Yelp heavily relies on Amazon S3 (Simple Storage Service) to store a wide variety of data, from images, logs, database backups, and more. Since data is stored on the cloud, we need to carefully manage how th…Mistral AIMistral AI raises 1.7B€ to accelerate technological progress with AIEtsyBuilding Etsy Buyer Profiles with LLMsEvery day, shoppers from Etsy's community of nearly 90M buyers visit our marketplace to search for unique, handmade, and vintage items. But with over 100 million listings, how do we help each buyer find exactly what they…Mistral AIMake Memory work for you.Mistral AIMake Memory work for you.Mistral AILe Chat. Custom MCP connectors. Memories.BAIR (Berkeley)What exactly does word2vec learn?What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language modeling task. Despite the fact that word2vec is a well-known prec…MozillaCRLite: Fast, private, and comprehensive certificate revocation checking in FirefoxFirefox is now the first and the only browser to deploy fast and comprehensive certificate revocation checking that does not reveal your browsing activity to anyone (not even to Mozilla). Tens of millions of TLS server c…EtsyContext engineering case studies: Etsy-specific question answeringThis post investigates the benefits and limitations of prompt engineering in two instances of AI-assisted onboarding relying on large language model (LLM) technology. Of particular interest is how truthful (and therefore…YelpExploring CHAOS: Building a Backend for Server-Driven UIA little while ago, we published a blog post on CHAOS: Yelp’s Unified Framework for Server-Driven UI. We strongly recommend reading that post first to gain a solid understanding of SDUI and the goals of CHAOS. This post…BAIR (Berkeley)Whole-Body Conditioned Egocentric Video Prediction.modal { display: none; position: fixed; z-index: 9999; padding-top: 50px; left: 0; top: 0; width: 100%; height: 100%; overflow: auto; background-color: rgba(0,0,0,0.9); } .modal-content { margin: auto; display: block; m…YelpRevenue Automation Series: Testing an Integration with Third-Party SystemBackground As described in the second blog post of Revenue Automation series, Revenue Data Pipeline processes a large amount of data via complex logic transformations to recognize revenue. Thus, developing a robust produ…PayPalPayPal Releases Agentic Toolkit to Accelerate CommerceThe following is a repost from the PayPal Developer Blog . Building on the release of PayPal’s MCP servers , PayPal is excited to introduce the PayPal Agentic Toolkit *. This toolkit empowers developers to seamlessly int…BAIR (Berkeley)Defending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign)Recent advances in Large Language Models (LLMs) enable exciting LLM-integrated applications. However, as LLMs have improved, so have the attacks against them. Prompt injection attack is listed as the #1 threat by OWASP t…BAIR (Berkeley)Repurposing Protein Folding Models for Generation with Latent DiffusionPLAID is a multimodal generative model that simultaneously generates protein 1D sequence and 3D structure, by learning the latent space of protein folding models. The awarding of the 2024 Nobel Prize to AlphaFold2 marks…PayPalPayPal Begins Rollout of MCP Servers to Accelerate Agentic CommerceThe following is a repost from the PayPal Developer Blog by Prakhar Mehrotra, SVP of Artificial Intelligence, PayPal At PayPal, we strive to make it easier for developers to access our services. Today, we are taking the…EtsyBehind the Scenes - A Glimpse into Tax CalculationsIn the past, sellers were responsible for managing and fulfilling their own tax obligations. However, more and more jurisdictions are now requiring marketplaces such as Etsy to collect the tax from buyers and remit the t…MozillaImproving Firefox Stability in the Enterprise by Reducing DLL InjectionBeginning in version 138, Firefox will offer an alternative to DLL injection for Data Loss Prevention (DLP) deployments in enterprise environments. DLL Injection DLL injection into Firefox is a topic we’ve covered on the…BAIR (Berkeley)Scaling Up Reinforcement Learning for Traffic Smoothing: A 100-AV Highway DeploymentTraining Diffusion Models with Reinforcement Learning We deployed 100 reinforcement learning (RL)-controlled cars into rush-hour highway traffic to smooth congestion and reduce fuel consumption for everyone. Our goal is…PayPalEstimating Incremental Lift in Customer Value (Delta CV) using Synthetic ControlHow we measure the impact of user actions and product adoptions at PayPal In today’s competitive digital landscape, understanding user interactions with your products is essential for driving revenue and building lasting…MozillaLaunching Interop 2025Interop 2025 continues the mission to make the web more consistent across browsers, building on 2024’s 95% interoperability score. This year, 19 focus areas target key developer needs and long-standing issues, including…EtsyAdopting Jetpack Compose for Etsy’s Android AppOne of our Guiding Principles at Etsy is that we “commit to our craft.” This means that we have a culture of learning, in which we’re constantly looking for opportunities to improve and learn, adopt industry best practic…MozillaIntroducing Uniffi for React Native: Rust-Powered Turbo ModulesMozilla and Filament have introduced Uniffi for React Native, a tool that allows developers to leverage the safety and performance benefits of Rust in cross-platform React Native apps. The post Introducing Uniffi for Rea…MozillaLlamafile v0.8.14: a new UI, performance gains, and moreDiscover the latest release of Llamafile 0.8.14, an open-source AI tool by Mozilla Builders. With a new command-line chat interface, enhanced performance, and support for powerful models, Llamafile makes it easy to run l…Mozilla0Din: A GenAI Bug Bounty Program – Securing Tomorrow’s AI TogetherAs AI continues to evolve, so do the threats against it. As these GenAI systems become more sophisticated and widely adopted, ensuring their security and ethical use becomes paramount. 0Din is a groundbreaking GenAI bug…MozillaAnnouncing Official Puppeteer Support for FirefoxWe’re pleased to announce that, as of version 23, the Puppeteer browser automation library now has first-class support for Firefox. This means that it’s now easy to write automation and perform end-to-end testing using P…EtsyMachine Learning in Content Moderation at EtsyAt Etsy, we’re focused on elevating the best of our marketplace to help creative entrepreneurs grow their businesses. We continue to invest in making Etsy a safe and trusted place to shop, so sellers’ extraordinary items…MozillaSnapshots for IPC FuzzingProcess separation remains one of the most important parts of the Firefox security model and securing our IPC (Inter-Process Communication) interfaces is crucial to keep privileges in the different processes separated. W…MozillaSponsoring sqlite-vec to enable more powerful Local AI applicationsToday we’re proud to announce the next Mozilla Builders project: sqlite-vec. Led by independent developer Alex Garcia, this project brings vector search functionality to the beloved SQLite embedded database. Alex has bee…EtsyEnhancing Cloud Usage Forecasting, Monitoring & OptimizingIn 2020, Etsy concluded its migration from an on-premise data center to the Google Cloud Platform (GCP). During this transition, a dedicated team of program managers ensured the migration's success. Post-migration, this…EtsyEfficient Visual Representation Learning And EvaluationEtsy features a diverse marketplace of unique handmade and vintage items. It’s a visually diverse marketplace as well, and computer vision has become increasingly important to Etsy as a way of enhancing our users’ shoppi…MozillaExperimenting with local alt text generation in Firefox NightlyFirefox 130 will introduce an experimental new capability to automatically generate alt-text for images using a fully private on-device AI model. The feature will be available as part of Firefox’s built-in PDF editor, an…PayPalScaling PayPal’s AI Capabilities with PayPal Cosmos.AI PlatformBy Jun Yang , Zhenyin Yang , and Srinivasan Manoharan , based on the AI/ML modernization journey taken by the PayPal Cosmos.AI Platform team in the past three years. Source: Dall-E 3 AI is a transformative technology tha…MozillaLlamafile’s progress, four months inWhen Mozilla’s Innovation group first launched the llamafile project late last year, we were thrilled by the immediate positive response from open source AI developers. It’s become one of Mozilla’s top three most-favorit…MozillaPorting a cross-platform GUI application to RustIn this blog post, we delve into the motivations for choosing Rust for our crash reporter, outline the unique challenges of designing an application that operates when the main browser has failed, and discuss the new arc…MozillaPrototype even faster with the Gradio UI for Figma component libraryIn the fast-paced world of generative AI, staying ahead means moving swiftly and smartly. That's why we've embraced Gradio, the low-code prototyping toolkit from Hugging Face, as our go-to for bringing new ideas to life.…GoogleGenerative AI to quantify uncertainty in weather forecastingPosted by Lizao (Larry) Li, Software Engineer, and Rob Carver, Research Scientist, Google Research Accurate weather forecasts can have a direct impact on people’s lives, from helping make routine decisions, like what to…GoogleAutoBNN: Probabilistic time series forecasting with compositional bayesian neural networksPosted by Urs Köster, Software Engineer, Google Research Time series problems are ubiquitous, from forecasting weather and traffic patterns to understanding economic trends. Bayesian approaches start with an assumption a…GoogleUsing AI to expand global access to reliable flood forecastsPosted by Yossi Matias, VP Engineering Research, and Grey Nearing, Research Scientist, Google Research Floods are the most common natural disaster , and are responsible for roughly $50 billion in annual financial damages…GoogleComputer-aided diagnosis for lung cancer screeningPosted by Atilla Kiraly, Software Engineer, and Rory Pilgrim, Product Manager, Google Research Lung cancer is the leading cause of cancer-related deaths globally with 1.8 million deaths reported in 2020. Late diagnosis d…GoogleSCIN: A new resource for representative dermatology imagesPosted by Pooja Rao, Research Scientist, Google Research Health datasets play a crucial role in research and medical education, but it can be challenging to create a dataset that represents the real world. For example, d…GoogleScreenAI: A visual language model for UI and visually-situated language understandingPosted by Srinivas Sunkara and Gilles Baechler, Software Engineers, Google Research Screen user interfaces (UIs) and infographics, such as charts, diagrams and tables, play important roles in human communication and huma…GoogleMELON: Reconstructing 3D objects from images with unknown posesPosted by Mark Matthews, Senior Software Engineer, and Dmitry Lagun, Research Scientist, Google Research A person's prior experience and understanding of the world generally enables them to easily infer what an object lo…EtsyMacramé: Untangling the Knot on the Etsy Android Listing ScreenEasily the most important and complex screen in the Buy on Etsy Android app is the listing screen, where all key information about an item for sale in the Etsy marketplace is displayed to buyers. Far from just a title an…GoogleHEAL: A framework for health equity assessment of machine learning performancePosted by Mike Schaekermann, Research Scientist, Google Research, and Ivor Horn, Chief Health Equity Officer Director, Google Core Health equity is a major societal concern worldwide with disparities having many causes.…GoogleCappy: Outperforming and boosting large multi-task language models with a small scorerPosted by Yun Zhu and Lijuan Liu, Software Engineers, Google Research Large language model (LLM) advancements have led to a new paradigm that unifies various natural language processing (NLP) tasks within an instruction-…GoogleTalk like a graph: Encoding graphs for large language modelsPosted by Bahare Fatemi and Bryan Perozzi, Research Scientists, Google Research Imagine all the things around you — your friends, tools in your kitchen, or even the parts of your bike. They are all connected in different…GoogleChain-of-table: Evolving tables in the reasoning chain for table understandingPosted by Zilong Wang, Student Researcher, and Chen-Yu Lee, Research Scientist, Cloud AI Team People use tables every day to organize and interpret complex information in a structured, easily accessible format. Due to th…GoogleHealth-specific embedding tools for dermatology and pathologyPosted by Dave Steiner, Clinical Research Scientist, Google Health, and Rory Pilgrim, Product Manager, Google Research There’s a worldwide shortage of access to medical imaging expert interpretation across specialties in…GoogleSocial learning: Collaborative learning with large language modelsPosted by Amirkeivan Mohtashami, Research Intern, and Florian Hartmann, Software Engineer, Google Research Large language models (LLMs) have significantly improved the state of the art for solving tasks specified using n…GoogleCroissant: a metadata format for ML-ready datasetsPosted by Omar Benjelloun, Software Engineer, Google Research, and Peter Mattson, Software Engineer, Google Core ML and President, MLCommons Association Machine learning (ML) practitioners looking to reuse existing datas…EtsyHow We Built The Deals Tab in Swift UIBalancing Engineering Ambition with Product Realism Introduction In July of 2023, Etsy’s App Updates team, responsible for the Updates feed in Etsy’s mobile apps, set off with an ambitious goal: to revamp the Updates tab…GoogleGoogle at APS 2024Posted by Kate Weber and Shannon Leon, Google Research, Quantum AI Team Today the 2024 March Meeting of the American Physical Society (APS) kicks off in Minneapolis, MN. A premier conference on topics ranging across phys…GoogleVideoPrism: A foundational visual encoder for video understandingPosted by Long Zhao, Senior Research Scientist, and Ting Liu, Senior Staff Software Engineer, Google Research An astounding number of videos are available on the Web, covering a variety of content from everyday moments p…PayPalLeveraging Spark 3 and NVIDIA’s GPUs to Reduce Cloud Cost by up to 70% for Big Data PipelinesBy Ilay Chen and Tomer Akirav At PayPal, hundreds of thousands of Apache Spark jobs run on an hourly basis, processing petabytes of data and requiring a high volume of resources. To handle the growth of machine learning…GoogleAdvances in private training for production on-device language modelsPosted by Zheng Xu, Research Scientist, and Yanxiang Zhang, Software Engineer, Google Language models (LMs) trained to predict the next word given input text are the key technology for many applications [ 1 , 2 ]. In Gbo…GoogleLearning the importance of training data under concept driftPosted by Nishant Jain, Pre-doctoral Researcher, and Pradeep Shenoy, Research Scientist, Google Research The constantly changing nature of the world around us poses a significant challenge for the development of AI model…GoogleDP-Auditorium: A flexible library for auditing differential privacyPosted by Mónica Ribero Díaz, Research Scientist, Google Research Differential privacy (DP) is a property of randomized mechanisms that limit the influence of any individual user’s information while processing and analyz…GoogleGraph neural networks in TensorFlowPosted by Dustin Zelle, Software Engineer, Google Research, and Arno Eigenwillig, Software Engineer, CoreML Objects and their relationships are ubiquitous in the world around us, and relationships can be as important to…GoogleIntervening on early readouts for mitigating spurious features and simplicity biasPosted by Rishabh Tiwari, Pre-doctoral Researcher, and Pradeep Shenoy, Research Scientist, Google Research Machine learning models in the real world are often trained on limited data that may contain unintended statistic…GoogleA decoder-only foundation model for time-series forecastingPosted by Rajat Sen and Yichen Zhou, Google Research Time-series forecasting is ubiquitous in various domains, such as retail, finance, manufacturing, healthcare and natural sciences. In retail use cases, for example, it…GoogleMobileDiffusion: Rapid text-to-image generation on-devicePosted by Yang Zhao, Senior Software Engineer, and Tingbo Hou, Senior Staff Software Engineer, Core ML Text-to-image diffusion models have shown exceptional capabilities in generating high-quality images from text prompt…GoogleMixed-input matrix multiplication performance optimizationsPosted by Manish Gupta, Staff Software Engineer, Google Research AI-driven technologies are weaving themselves into the fabric of our daily routines, with the potential to enhance our access to knowledge and boost our ov…GoogleExphormer: Scaling transformers for graph-structured dataPosted by Ameya Velingker, Research Scientist, Google Research, and Balaji Venkatachalam, Software Engineer, Google Graphs , in which objects and their relations are represented as nodes (or vertices) and edges (or links…PayPalDeclarative Feature Engineering at PayPalPhoto by fabio on Unsplash PayPal supports over 400 million active consumers and merchants worldwide. Every minute there are several thousand payment transactions. To prevent fraud in real-time at such a scale, we need t…PayPalStreamlining Developer Productivity with the PayPal Visual Studio Code ExtensionIn the ever-evolving landscape of software development, productivity and efficiency have become paramount to success. Developers are constantly juggling multiple tasks, from navigating complex codebases to integrating th…PayPalManaging Recurring Payments with Apple Pay Using PayPalRecurring payments have become an integral part of the modern digital economy, offering convenience and predictability for both consumers and businesses. Our previous post highlighted different methods of integrating App…PayPalAccept E-Commerce Payments Easily with PayPal’s Buttons ComponentAccepting online payments is now a universal must-have, catering to everyone from solo entrepreneurs to massive global corporations. PayPal’s Standard Checkout allows for seamless integration of PayPal’s Payment Buttons…PayPalWhy You Should Attend PayPal’s Developer Meetup at Money20/20The world of technology is constantly evolving, and developers are at the forefront of this dynamic landscape. Staying updated on the latest trends, tools, and innovations is not just a choice but a necessity for those i…EtsyThe AR Measuring Box: Etsy&#39;s answer to Big Tape MeasureA little while ago, Etsy introduced a new feature in its iOS app that could place Etsy sellers' artwork on a user's wall using Apple's Augmented Reality (AR) tools. It let them visualize how a piece would look in their s…EtsyThe So-fine Real-time ML ParadigmIntroduction Each year, Etsy hosts an event known as “CodeMosaic” - an internal hackathon in which Etsy admin propose and build bold advances quickly in our technology across a number of different themes. People across E…EtsyLeveraging Real-Time User Actions to Personalize Etsy AdsIntroduction Personalization is vital to connect our unique marketplace to the right buyer at the right time. Etsy has recently introduced a novel, general approach to personalizing ML models based on encoding and learni…Stanford AI LabLinkBERT: Improving Language Model Training with Document LinkLanguage Model Pretraining Language models (LMs), like BERT 1 and the GPT series 2 , achieve remarkable performance on many natural language processing (NLP) tasks. They are now the foundation of today’s NLP systems. 3 T…Stanford AI LabStanford AI Lab Papers and Talks at ACL 2022The 60th Annual Meeting of the Association for Computational Linguistics (ACL) 2022 is taking place May 22nd - May 27th. We’re excited to share all the work from SAIL that’s being presented, and you’ll find links to pape…Stanford AI LabStanford AI Lab Papers and Talks at ICLR 2022The International Conference on Learning Representations (ICLR) 2022 is being hosted virtually from April 25th - April 29th. We’re excited to share all the work from SAIL that’s being presented, and you’ll find links to…Stanford AI LabDiscovering the systematic errors made by machine learning modelsDiscovering systematic errors with cross-modal embeddings In this blog post, we introduce Domino, a new approach for discovering systematic errors made by machine learning models. We also discuss a framework for quantita…Stanford AI LabGrading Complex Interactive Coding Programs with Reinforcement Learning[Summary] tl;dr: A tremendous amount of effort has been poured into training AI algorithms to competitively play games that computers have traditionally had trouble with, such as the retro games published by Atari, Go, D…Stanford AI LabUnderstanding Deep Learning Algorithms that Leverage Unlabeled Data, Part 1: Self-trainingDeep models require a lot of training examples, but labeled data is difficult to obtain. This motivates an important line of research on leveraging unlabeled data, which is often more readily available. For example, larg…Stanford AI LabStanford AI Lab Papers and Talks at AAAI 2022The 36th AAAI Conference on Artificial Intelligence (AAAI 2022) is being hosted virtually from February 22th - March 1st. We’re excited to share all the work from SAIL that’s being presented, and you’ll find links to pap…Stanford AI LabHow to Improve User Experience (and Behavior): Three Papers from Stanford's Alexa Prize TeamIntroduction In 2019, Stanford entered the Alexa Prize Socialbot Grand Challenge 3 for the first time, with its bot Chirpy Cardinal , which went on to win 2nd place in the competition. In our previous post , we discussed…Stanford AI LabReward Isn't Free: Supervising Robot Learning with Language and Video from the WebThis work was conducted as part of SAIL and CRFM . Deep learning has enabled improvements in the capabilities of robots on a range of problems such as grasping 1 and locomotion 2 in recent years. However, building the qu…Stanford AI LabBanditPAM: Almost Linear-Time k-medoids Clustering via Multi-Armed BanditsTL;DR Want something better than \(k\)-means? Our state-of-the-art \(k\)-medoids algorithm from NeurIPS, BanditPAM, is now publicly available! \(\texttt{pip install banditpam}\) and you're good to go! Like the \(k\)-mean…Stanford AI LabStanford AI Lab Papers and Talks at NeurIPS 2021The thirty-fifth Conference on Neural Information Processing Systems (NeurIPS) 2021 is being hosted virtually from Dec 6th - 14th. We’re excited to share all the work from SAIL that’s being presented at the main conferen…Stanford AI LabStanford AI Lab Papers at CoRL 2021The Conference on Robot Learning (CoRL 2021) will take place next week. We’re excited to share all the work from SAIL that will be presented, and you’ll find links to papers, videos and blogs below. Feel free to reach ou…Stanford AI LabStanford AI Lab Papers at EMNLP/CoNLL 2021The 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021) will take place next week, colocated with CoNLL 2021. We’re excited to share all the work from SAIL that will be presented, and you’ll…Stanford AI LabSelective Classification Can Magnify Disparities Across GroupsSelective classification, where models are allowed to “abstain” when they are uncertain about a prediction, is a useful approach for deploying models in settings where errors are costly. For example, in medicine, model e…Stanford AI LabStanford AI Lab Papers at ICCV 2021The International Conference on Computer Vision (ICCV 2021) will be hosted virtually next week. We’re excited to share all the work from SAIL that will be presented, and you’ll find links to papers, videos and blogs belo…