Jocelyn

We curate AI news developments for you.

Strong negative catalystStrong positive catalyst

Google Research BlogResearch

Toward provably private learning from federated data

Google Research introduced a new federated learning system using Trusted Execution Environments (TEEs) to provide verifiable privacy guarantees, improving accuracy and speed. Gboard has adopted the system, achieving faster compute times with encrypted data processed only within TEEs.

GOOGLModest positive

CoreWeave BlogResearch

Why I Built Fugue: Controlled Experiments for Agent Systems

CoreWeave's Fugue tool was developed to enable controlled experiments for agent systems, focusing on improving task execution through structured prompts and tool integration. A 77% success rate was observed with a thinner API-based runner compared to a full command-line harness.

Cohere BlogResearch

RCP-nDCG@10: A New Approach to Enterprise Retrieval Quality | Cohere

Cohere introduced RCP-nDCG@10, a new retrieval quality metric using a calibrated AI judge to assess relevance against explicit criteria. It addresses limitations of nDCG by evaluating all retrieved documents, not just labeled ones, and was validated with a 0.91 AUC in human relevance prediction.

Cohere BlogResearch

AI Future of Work Evidence Gap | Cohere Labs

A 2023 paper estimated 80% of U.S. workers have tasks exposed to large language models, cited by the IMF and U.S. Senate, but its 2023 model and American taxonomy limit its current applicability. Newer tools aim to provide more dynamic, representative evidence for policy decisions.

Cohere BlogResearch

Cohere Blog | Research

Cohere launched Cohere Transcribe, a new open-source speech recognition tool. The product aims to set a new standard in speech recognition technology.

Hugging Face BlogResearch

AutoSynthData: Generating Training Data for Enterprise Agents

Hugging Face's ServiceNow CoreAI developed AutoSynthData to generate training data for enterprise agents by identifying capability gaps and creating tasks that test those weaknesses. The system uses a stronger teacher model to validate tasks, ensuring they are feasible, realistic, and difficult for the target model.

NOWModest positive

Apple Machine Learning ResearchResearch

Limits of Confidence in Diffusion

Apple Machine Learning Research published a paper titled "Limits of Confidence in Diffusion," highlighting that discrete diffusion methods generate sequences with dependencies between tokens. The study found that generated distributions on a synthetic task were 29× the sampling-noise floor total variation.

Amazon ScienceResearch

Graph-centric agentic intelligence

Amazon developed graph-centric agentic intelligence to create a "digital twin" for networks, enabling faster failure isolation. The system uses cascaded graph analytics to identify root causes in minutes, demonstrated with NTT DOCOMO achieving 95% accuracy in 30 seconds.

AMZNModest positive

Apple Machine Learning ResearchResearch

RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback

Apple researchers introduced RLTL;DR, a reinforcement learning method that uses self-generated feedback to improve performance. On challenging coding tasks, it achieved 12–13% pass rate without insights during evaluation, outperforming a Qwen 3.5 9B model.

AAPLModest positive

Google DeepMind BlogResearch

Introducing SynthID Bio

Google DeepMind introduced SynthID Bio, a watermarking technology for AI-generated proteins that preserves biological function. In tests, watermarked proteins matched unwatermarked versions in binding affinity and diversity across three targets: VEGF-A, SARS-CoV-2 spike protein RBD, and PD-L1.

GOOGLModest positive

Apple Machine Learning ResearchResearch

SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation

Apple Machine Learning Research introduced SCLATE, a substrate for continual-learning agent training and evaluation. SCLATE enabled testing ten unmodified agent configurations, showing Qwen3.5-4B improved SWE-bench pass rate by 16.7 points and reduced file lines read by 6.8×.

AAPLModest positive

Apple Machine Learning ResearchResearch

On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study

Apple Machine Learning Research studied the trade-off between effectiveness and fluency in LLM conditioning, finding that efficient methods often reduce fluency. Activation steering was less effective on instruction-tuned models, while prompting and fine-tuning worked better for injection but not removal.

Apple Machine Learning ResearchResearch

Faster Rates for Federated Variational Inequalities

Apple researchers introduced LIPPAX, a new algorithm for federated variational inequalities, improving convergence rates and reducing client drift. The study, conducted while at Apple, builds on Local Extra SGD and applies to bounded Hessian, operator, and low-variance settings.

AAPLModest positive

Amazon ScienceResearch

A kernel-centric path to real-time video generation on Trainium

Reactor and Amazon Neuron Science optimized real-time video generation using the Neuron Kernel Interface on Trainium, focusing on autoregressive diffusion models. They achieved consistent 30-second video generation with Rolling Forcing, addressing memory and latency challenges.

AMZNModest positive