logo
|
Blog
  • Yetter
  • OwLite
  • Fits on Chips
  • SqueezeBits
  • ENKR

Unlock the Potential of AI

Deploy your AI with Maximal Efficiency
Yeonjoon Jung's avatar
Yeonjoon Jung
Reliable & Scalable Synthetic Data for Physical AI (Part 2): Making Cosmos 3.1 x Faster for Production

Reliable & Scalable Synthetic Data for Physical AI (Part 2): Making Cosmos 3.1 x Faster for Production

Explore why Physical AI deployment needs synthetic data at scale with Squeezebits' research and discover how to overcome inference bottlenecks to accelerate Roboost Agent.
Jongho Lee's avatar
Daehyun Ahn's avatar
Yeonjoon Jung's avatar
Semin Kim's avatar
Seungryeol Kim's avatar
Mar 11, 2026
Tech Insight
Reliable & Scalable Synthetic Data for Physical AI (Part 1): Taming NVIDIA Cosmos with RoBoost Agent

Reliable & Scalable Synthetic Data for Physical AI (Part 1): Taming NVIDIA Cosmos with RoBoost Agent

Scaling Physical AI requires reliable synthetic data. Learn how RoBoost Agent integrates NVIDIA Cosmos to transform world models into trustworthy data engines for robotics and autonomous driving.
Daehyun Ahn's avatar
Jongho Lee's avatar
Yeonjoon Jung's avatar
Semin Kim's avatar
Seungryeol Kim's avatar
Feb 25, 2026
Tech Insight
Winning both speed and quality: How Yetter deals with diffusion models

Winning both speed and quality: How Yetter deals with diffusion models

Explore how the Yetter Inference Engine overcomes the limitations of step caching and model distillation for diffusion models. We analyze latency, diversity, quality, and negative-prompt handling to reveal what truly matters for scalable, real-time image generation.
Yeonjoon Jung's avatar
Oct 31, 2025
Product
GraLoRA: Boosting Fine-Tuning Accuracy Without Extra Cost

GraLoRA: Boosting Fine-Tuning Accuracy Without Extra Cost

LoRA excels at efficient fine-tuning but suffers at higher ranks due to gradient entanglement. We introduce GraLoRA, which addresses these issues through finer-grained, block-wise updates, significantly enhancing performance and expressivity without overhead. GraLoRA outperforms LoRA across tasks, achieving up to +8.5% improvement in HumanEval+ Pass@1.
Yeonjoon Jung's avatar
Jul 21, 2025
Tech Insight
[vLLM vs TensorRT-LLM] #13. Vision-Language Models

[vLLM vs TensorRT-LLM] #13. Vision-Language Models

This article provides a comparative analysis of serving vision-language models on vLLM and TensorRT-LLM.
Yeonjoon Jung's avatar
Jan 20, 2025
Tech Insight
[vLLM vs TensorRT-LLM] #12. Automatic Prefix Caching

[vLLM vs TensorRT-LLM] #12. Automatic Prefix Caching

This article provides a comparative analysis of automatic prefix caching.
Daehyun Ahn's avatar
Yeonjoon Jung's avatar
Taesu Kim's avatar
Huijong Jeong's avatar
Dec 23, 2024
Tech Insight
[vLLM vs TensorRT-LLM] #11. Speculative Decoding

[vLLM vs TensorRT-LLM] #11. Speculative Decoding

This article provides a comparative analysis of speculative decoding.
Daehyun Ahn's avatar
Yeonjoon Jung's avatar
Dec 09, 2024
Tech Insight
[vLLM vs TensorRT-LLM] #2. Towards Optimal Batching for LLM Serving

[vLLM vs TensorRT-LLM] #2. Towards Optimal Batching for LLM Serving

This article provides a comparative analysis of vLLM and TensorRT-LLM frameworks, focusing on batching configurations and thoroughly examining the effects of maximum batch size and maximum number of tokens.
Yeonjoon Jung's avatar
Oct 11, 2024
Tech Insight
[vLLM vs TensorRT-LLM] #1. An Overall Evaluation

[vLLM vs TensorRT-LLM] #1. An Overall Evaluation

This article provides a comparative analysis of vLLM and TensorRT-LLM frameworks for serving LLMs, evaluating their performance based on key metrics like throughput, TTFT, and TPOT to offer insights for practitioners in optimizing LLM deployment strategies.
Yeonjoon Jung's avatar
Oct 01, 2024
Tech Insight

The official SqueezeBits Tech blog

RSS·Powered by Inblog