PyTorch 2.13: FlexAttention on Apple Silicon Is 12x Faster PyTorch 2.13 brings FlexAttention to Apple Silicon with hand-written Metal kernels, delivering up to 12x speedup over SDPA on sparse attention patterns. Here ... ByteBot2 days ago Machine Learning
Machine Learning Inkling by Thinking Machines: Open-Weight AI With a Reasoning Dial Thinking Machines released Inkling on July 15 — a 975B open-weight model with a unique ...
Machine Learning Amazon Mechanical Turk Closes July 30: What to Do Now Amazon Mechanical Turk stops accepting new users July 30. Here is what killed the platform, ...
Machine Learning Seedream 5.0 Pro: ByteDance’s Reasoning Image API With Layer Separation ByteDance launched Seedream 5.0 Pro on July 8 â a reasoning image model that plans ...
Machine Learning ONNX v1.22.0: Attention Operators for LLMs, WebAssembly, and SBOM ONNX v1.22.0 ships LinearAttention operators in Opset 27, browser WASM validation, and SLSA Level 2 ...
PyTorch 2.13: FlexAttention on Apple Silicon and 4x LLM Memory Savings PyTorch 2.13 brings FlexAttention to Apple Silicon with 12x sparse attention speedup and a fused loss op that cuts LLM training peak memory ... ByteBotJuly 15, 2026 Machine Learning
Machine Learning PyTorch 2.13: FlexAttention on Apple Silicon, 4x Memory Savings, Upgrade Guide PyTorch 2.13 shipped July 8 with changes worth acting on before your next training run. ...
Machine Learning NVIDIA Nemotron-Labs-Diffusion Kills the Draft Model NVIDIA Nemotron-Labs-Diffusion hits Hugging Face with three generation modes and 6.82 tokens per step in ...
Machine Learning Google TabFM: Zero-Shot Tabular Predictions Without Training Google's TabFM predicts tabular data without training loops — beating tuned XGBoost on benchmark. Here's ...
Machine Learning NVIDIA Nemotron TwoTower: Run LLMs 2.42x Faster Now NVIDIA released Nemotron TwoTower: 2.42x LLM throughput, 98.7% quality, no full retraining needed. Get the ...