New STAC Research report: benchmarking the infrastructure for agentic quantitative research. Given a goal, agentic systems will propose an approach, write the code, train a model, evaluate it, and decide what to try next. Using a workload built from a recorded research session contributed by a large US market maker, our latest report leverages the SwarmOne simulator and toolkit to measure a Lambda 1-Click Cluster of NVIDIA B200 GPUs at four levels of load and shows the performance of the platform as more agents run on the cluster. STAC is forming a working group under the STAC Benchmark Council to develop a rigorous benchmark for this class of workload. If you're interested in helping shape it, contact us at https://lnkd.in/dd-MEsfa Read the report: https://lnkd.in/dBzxzYUZ Subscribe to STAC Insights to read the accompanying STAC Configuration Disclosure with exact product versions and tuning detail: https://lnkd.in/d6S3vAXB
STAC Research Report: Agentic Quantitative Research Infrastructure Benchmarking
More Relevant Posts
-
STAC Research just benchmarked the infrastructure behind agentic quantitative research. The workload comes from a real research session, contributed by a large US market maker, and was replayed with SwarmSim on a Lambda 1-Click Cluster of NVIDIA B200 GPUs at four levels of load. The question is a practical one: what happens to the platform as more agents run on the cluster. SwarmOne is glad to see SwarmSim used as the simulator for this work. This is the kind of measurement the industry has been missing.
New STAC Research report: benchmarking the infrastructure for agentic quantitative research. Given a goal, agentic systems will propose an approach, write the code, train a model, evaluate it, and decide what to try next. Using a workload built from a recorded research session contributed by a large US market maker, our latest report leverages the SwarmOne simulator and toolkit to measure a Lambda 1-Click Cluster of NVIDIA B200 GPUs at four levels of load and shows the performance of the platform as more agents run on the cluster. STAC is forming a working group under the STAC Benchmark Council to develop a rigorous benchmark for this class of workload. If you're interested in helping shape it, contact us at https://lnkd.in/dd-MEsfa Read the report: https://lnkd.in/dBzxzYUZ Subscribe to STAC Insights to read the accompanying STAC Configuration Disclosure with exact product versions and tuning detail: https://lnkd.in/d6S3vAXB
To view or add a comment, sign in
-
-
Kevin Matcham will be talking about how Bentaus is flexing live GPU workloads using AMD Instinct MI325X and MI355X GPUs through Ziani Systems . We’ve now validated GPU power modulation across multiple AMD GPU platforms, with some amazing results around response time, power control and workload performance. We’ll be publishing our GPU Flexibility Validation Report shortly with the results and methodology. Kevin will be giving a presenting some of the results today. Excited to share more at AI Infra Summit and the AMD booth. Join us at 4pm CT Today! AI is becoming a flexible energy resource and we are at the forefront of it with our amazing partners.
To view or add a comment, sign in
-
We’ve now validated GPU power modulation across multiple AMD GPU platforms, with some amazing results around response time, power control and workload performance. We’ll be publishing our GPU Flexibility Validation Report shortly with the results and methodology. Kevin Matcham will be giving a presenting some of the results today.
Kevin Matcham will be talking about how Bentaus is flexing live GPU workloads using AMD Instinct MI325X and MI355X GPUs through Ziani Systems . We’ve now validated GPU power modulation across multiple AMD GPU platforms, with some amazing results around response time, power control and workload performance. We’ll be publishing our GPU Flexibility Validation Report shortly with the results and methodology. Kevin will be giving a presenting some of the results today. Excited to share more at AI Infra Summit and the AMD booth. Join us at 4pm CT Today! AI is becoming a flexible energy resource and we are at the forefront of it with our amazing partners.
To view or add a comment, sign in
-
NVIDIA Dynamo fits one layer above the inference engine. SGLang, vLLM, and TensorRT-LLM still do the engine work; Dynamo is about coordinating that workload across GPUs and nodes. That separation matters when a team already has an engine in production, because scaling is not the same problem as serving a model on one machine. The five-minute walkthrough by Vishakha Sadhwani makes that boundary easy to see. #InferenceServing #GPUInfrastructure
Already running an inference engine? So where does NVIDIA Dynamo fit in? In five minutes, we break down how Dynamo sits around engines like SGLang, vLLM and TensorRT-LLM to scale inference across GPUs and nodes. Full video by our very own Vishakha Sadhwani ➡️ https://lnkd.in/dK8HW4yU
To view or add a comment, sign in
-
Haocheng Xi and the OpenVDN team released Video Delta Net (VDN), a genuinely different approach to speeding up MiniMax H3 than the projects we've covered so far. Rather than reducing denoising steps or optimizing existing attention computation, VDN rebuilds the attention mechanism itself. Softmax attention accounts for more than 85% of H3's runtime, and its cost grows quadratically with clip length. VDN splits the workload: a sliding-window softmax branch handles nearby frames to preserve quality, while a frame-wise linear attention branch handles long-range context at a fraction of the cost. Both merge into H3's frozen backbone via LoRA adapters, no retraining of the original weights required. On 8x B200 GPUs, this produces a 14.4-second clip in 11.23 seconds, genuinely competitive with the fastest H3 variants covered here. The team's own claimed 75-90x speedup is measured against an unoptimized single-GPU baseline, worth knowing before treating that number as directly comparable to other projects' benchmarks. One real catch: the weights are licensed under terms that explicitly exclude the EU, UK, South Korea, and the United States from authorized use, a genuine constraint on who can actually deploy this legally. https://lnkd.in/gmHP6a8r
To view or add a comment, sign in
-
Cochl.Sense Nano is live! Cochl.Sense Nano is our smallest acoustic AI model yet. Compared to the Cochl.Sense Edge SDK, it is over 10x smaller in model size and over 50x smaller in memory footprint, with comparable recognition performance. That is what puts sound AI on hardware it could not reach before: • MCUs such as STM32, ESP32, and Arm Ethos-U55 • Low power NPUs such as Syntiant NDP120 • Battery powered and always-on devices where the power and memory budget were the blockers Small footprint, full capability. Let's talk about what Nano can run on next.
To view or add a comment, sign in
-
-
🌟 MFEM 4.10 released with many new features: ✓ GPU partial assembly for simplicial Bernstein H1 mass and diffusion ✓ FindPointsGSLIB point search and interpolation on surface meshes ✓ expanded GPU kernels and GPU-aware MPI communication ✓ parallel nonuniform anisotropic refinement on quad and hex meshes ✓ NVIDIA cuDSS sparse direct solver and MPI-distributed Ginkgo solvers ✓ GridFunction extrema estimates by piece-wise linear bounds ✓ complex-valued mixed sesquilinear forms ✓ positive-weight quadrature rules on triangles and tets ✓ Gmsh 4.1 mesh format support (ASCII and binary) ✓ 7 new examples and a new electrostatic Particle-In-Cell miniapp ✓ and much more! To download and for more information, visit https://mfem.org. We welcome any and all feedback at https://lnkd.in/exz_9xtP
To view or add a comment, sign in
-
-
What does it take to deliver high-quality results in seconds at massive scale? @PerplexityAI serves 50M+ queries daily across 15 models, including NVIDIA Nemotron 3 Ultra. That takes high-performance infrastructure and smarter routing without compromising quality, speed, or cost. That’s why Perplexity has built on NVIDIA generation after generation. With NVIDIA Blackwell, the larger NVLink domain enables models to scale across multiple nodes without sacrificing latency. Learn more: https://bit.ly/4hn8LkB
To view or add a comment, sign in
-
What does it take to deliver high-quality results in seconds at massive scale? @PerplexityAI serves 50M+ queries daily across 15 models, including NVIDIA Nemotron 3 Ultra. That takes high-performance infrastructure and smarter routing without compromising quality, speed, or cost. That’s why Perplexity has built on NVIDIA generation after generation. With NVIDIA Blackwell, the larger NVLink domain enables models to scale across multiple nodes without sacrificing latency. Learn more: https://bit.ly/4ivnpIF
To view or add a comment, sign in
-
What does it take to deliver high-quality results in seconds at massive scale? @PerplexityAI serves 50M+ queries daily across 15 models, including NVIDIA Nemotron 3 Ultra. That takes high-performance infrastructure and smarter routing without compromising quality, speed, or cost. That’s why Perplexity has built on NVIDIA generation after generation. With NVIDIA Blackwell, the larger NVLink domain enables models to scale across multiple nodes without sacrificing latency. Learn more: https://bit.ly/4ivnpIF
To view or add a comment, sign in