Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#gpu
Tag79건Article 79

#gpu

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#nvidia공동문서 57 · 연관도 30%#capex-cycle공동문서 42 · 연관도 29%#applications공동문서 79 · 연관도 27%#ai-architecture공동문서 48 · 연관도 24%#llm공동문서 77 · 연관도 24%#agent-routing공동문서 32 · 연관도 15%#service-design공동문서 24 · 연관도 14%#retrieval-index공동문서 19 · 연관도 13%#ai-infrastructure공동문서 20 · 연관도 13%#agent-memory공동문서 27 · 연관도 13%
Hotter Than a Hot Tub: The 45°C Breakthrough to Cool AI’s Biggest Machines
Article2026년 6월 22일

Hotter Than a Hot Tub: The 45°C Breakthrough to Cool AI’s Biggest Machines

NVIDIA는 Rubin 세대 AI 인프라에서 45°C까지 동작하는 100% 액체 냉각을 통해 팬, 냉각수탑 의존, 물 소비를 크게 줄이면서 대규모 AI 데이터센터의 에너지 효율을 높이려 한다.

Josh Parker
#nvidia#ai-architecture#ai-infrastructure#capex-cycle
From Materials Simulation to Experimental Astronomy, New NVIDIA AI Software Unlocks Scientific Discoveries
Article2026년 6월 22일

From Materials Simulation to Experimental Astronomy, New NVIDIA AI Software Unlocks Scientific Discoveries

NVIDIA는 ISC에서 cuPhoton, DAQIRI, ALCHEMI를 소개하며 천문 데이터 처리, 실시간 실험 데이터 분석, 화학·재료 시뮬레이션을 GPU 기반 파이프라인으로 크게 가속한다고 밝혔다.

Chris Porter
#nvidia#service-design#agent-memory#agent-routing
NAIRR Science Program Reshapes Scientific Research, Powered by NVIDIA AI Infrastructure
Article2026년 6월 22일

NAIRR Science Program Reshapes Scientific Research, Powered by NVIDIA AI Infrastructure

미국 NSF의 NAIRR 파일럿은 NVIDIA DGX 기반 AI 인프라와 기술 지원을 통해 700개 이상의 연구 프로젝트에서 물리 시뮬레이션, 에너지 소재 탐색, 감염병 감시 같은 과학 연구의 속도와 범위를 확장하고 있다.

Zoe Kessler
#nvidia#ai-architecture#capex-cycle#search-advertising
NVIDIA Vera CPU Opens the Way for Agentic Scientific AI at Los Alamos National Laboratory
Article2026년 6월 22일

NVIDIA Vera CPU Opens the Way for Agentic Scientific AI at Los Alamos National Laboratory

로스앨러모스 국립연구소는 HPE와 NVIDIA 기반의 Mission, Vision, Veritas 슈퍼컴퓨터에 Vera CPU를 도입해 과학 시뮬레이션과 에이전트형 AI 연구를 가속하려 한다.

Chris Porter
#nvidia#ai-architecture#ai-infrastructure#capex-cycle
We got local models to triage the OpenClaw repo for FREE!*
Article2026년 6월 22일

We got local models to triage the OpenClaw repo for FREE!*

기존 NVIDIA GB10 장비에서 Gemma와 Qwen을 제한된 에이전트 하네스로 구동해 OpenClaw 이슈·PR을 실시간 분류하고, 필요한 항목만 디스코드로 전달하는 로컬 트리아지 시스템을 구축·평가한 사례다.

huggingface.co
#openclaw#anthropic#ai-architecture#llm
From the Hugging Face Hub to robot hardware with Strands Agents and LeRobot
Article2026년 6월 17일

From the Hugging Face Hub to robot hardware with Strands Agents and LeRobot

이 글은 Strands Robots가 LeRobot 데이터셋, 시뮬레이션, 정책 실행, 하드웨어 경로를 하나의 에이전트 루프로 묶어 Hugging Face Hub의 데모 데이터에서 실제 로봇 실행까지 이어지는 흐름을 설명한다.

huggingface.co
#anthropic#service-design#agent-deployment#agent-routing
GLM-5.2: Built for Long-Horizon Tasks
Article2026년 6월 17일

GLM-5.2: Built for Long-Horizon Tasks

GLM 5.2는 안정적인 100만 토큰 문맥, 강화된 코딩 성능, 효율적인 장문맥 구조와 서빙·강화학습 체계를 결합해 장시간 수행되는 복잡한 엔지니어링 작업을 겨냥한 공개형 주력 모델이다.

huggingface.co
#service-design#ai-architecture#agent-routing#llm
Parallelize speculative decoding with P-EAGLE on Amazon SageMaker AI
Article2026년 6월 16일

Parallelize speculative decoding with P-EAGLE on Amazon SageMaker AI

P EAGLE은 EAGLE식 추측 디코딩의 순차적 드래프트 병목을 병렬 드래프팅으로 제거해 SageMaker JumpStart에서 지원 모델의 추론 처리량을 높이는 방법입니다.

@https://twitter.com/pymhq
#nvidia#service-design#ai-architecture#agent-deployment
Want to get a data center online quickly? Give it some flex.
Article2026년 6월 16일

Want to get a data center online quickly? Give it some flex.

데이터센터를 더 빨리 가동하기 위한 해법으로 새 발전소 건설만이 아니라, 전력 수요가 치솟는 순간 데이터센터의 소비전력을 낮추는 ‘전력 유연성’이 주목받고 있다.

technologyreview.com
#nvidia#privacy-design#service-design#ai-infrastructure
Deepfakes Are Free Now. Your Company’s Phone System Is Not.
Article2026년 6월 14일

Deepfakes Are Free Now. Your Company’s Phone System Is Not.

이 글은 무료·저가 도구로 음성 복제와 딥페이크가 쉬워진 상황에서, 기업의 전화·음성 채널이 가장 방치된 사회공학 공격면이 되고 있다고 경고한다.

medium.com
#privacy-design#llm#applications#gpu
The Hill-Climbing Machine. What Satya Nadella Got Right — and the…
Article2026년 6월 14일

The Hill-Climbing Machine. What Satya Nadella Got Right — and the…

이 글은 사티아 나델라가 말한 AI 시대의 ‘학습 루프’가 옳지만, 대부분의 중소기업에는 그 루프를 작동시킬 데이터 정리, 프로세스 문서화, 거버넌스 기반이 먼저 필요하다고 주장한다.

medium.com
#anthropic#privacy-design#service-design#ai-architecture
NVIDIA Confidential Computing to Help Expand Apple’s Private Cloud Compute
Article2026년 6월 9일

NVIDIA Confidential Computing to Help Expand Apple’s Private Cloud Compute

엔비디아 Confidential Computing GPU가 Apple Private Cloud Compute의 Google Cloud 확장에 쓰이며, Apple Intelligence의 일부 차세대 기능을 위한 서버 측 추론과 개인정보 보호를 지원한다.

Avinash Ahuja
#nvidia#apple-silicon#privacy-design#ai-architecture
DeepSeek-V4: a million-token context that agents can actually use
Article2026년 6월 8일

DeepSeek-V4: a million-token context that agents can actually use

DeepSeek V4는 최고 벤치마크 점수보다 100만 토큰 문맥을 실제 에이전트 작업에서 감당하게 만드는 긴 문맥 효율, 도구 호출 지속성, 샌드박스 기반 학습 인프라에 초점을 둔 모델이다.

huggingface.co
#ai-architecture#agent-memory#agent-routing#context-compression
An Interview with Microsoft CEO Satya Nadella About Finding Core Competencies – Stratechery by Ben Thompson
Article2026년 6월 4일

An Interview with Microsoft CEO Satya Nadella About Finding Core Competencies – Stratechery by Ben Thompson

사티아 나델라는 AI 전환기에서 Microsoft의 핵심 역량을 ‘모두가 자기만의 학습 기계를 만들 수 있게 하는 신뢰받는 플랫폼’으로 재정의하며, OpenAI와 자체 MAI 모델을 함께 활용하는 전략을 설명한다.

stratechery.com
#anthropic#hotel-review#token-efficiency#ai-safety
On the Shifting Global Compute Landscape
Article2026년 5월 26일

On the Shifting Global Compute Landscape

미국 중심이던 인공지능 연산 생태계가 수출 통제, 중국산 칩의 성장, 개방형 모델과 연산 효율 기술의 확산을 계기로 중국을 포함한 다극적 하드웨어·소프트웨어 구조로 재편되고 있다.

huggingface.co
#nvidia#linear-attention#ai-architecture#agent-memory
Learn the Hugging Face Kernel Hub in 5 Minutes
Article2026년 5월 18일

Learn the Hugging Face Kernel Hub in 5 Minutes

허깅페이스 커널 허브는 파이썬·파이토치·쿠다 환경에 맞는 사전 최적화 연산 커널을 허브에서 내려받아, 복잡한 로컬 빌드 없이 모델의 특정 연산에 적용하도록 돕는 체계다.

huggingface.co
#ai-architecture#agent-routing#llm#applications
Building the compute infrastructure for the Intelligence Age
Article2026년 4월 29일

Building the compute infrastructure for the Intelligence Age

OpenAI는 Stargate를 통해 급증하는 AI 수요에 대응할 대규모 컴퓨트 인프라를 파트너·지역사회와 함께 구축하고, 그 혜택을 더 넓게 확산하겠다고 설명한다.

openai.com
#openai#privacy-design#ai-safety#llm
Unweight: how we compressed an LLM 22% without sacrificing quality
Article2026년 4월 17일

Unweight: how we compressed an LLM 22% without sacrificing quality

Unweight는 LLM 가중치의 BF16 지수 바이트를 무손실로 압축하고 GPU 온칩 메모리에서 바로 복원해, 출력 품질을 유지하면서 모델 크기와 HBM 메모리 대역폭 부담을 줄이는 추론용 압축 시스템이다.

blog.cloudflare.com
#model-scaling#ai-architecture#agent-memory#context-compression
Building the foundation for running extra-large language models
Article2026년 4월 16일

Building the foundation for running extra-large language models

Cloudflare는 Workers AI에서 Kimi K2.5 같은 초대형 오픈소스 언어 모델을 빠르고 효율적으로 운영하기 위해 하드웨어 구성, prefill/decode 분리, 프롬프트 캐싱, KV 캐시 최적화, speculative decoding, 자체 추론 엔진 Infire를 조합하고 있다고 설명한다.

@_mchenco
#kimi-k2-5#nvidia#ai-architecture#agent-memory
The inevitable need for an open model consortium
Article2026년 4월 11일

The inevitable need for an open model consortium

저자는 훈련 비용과 수익 압박이 커질수록 근접 프런티어 수준의 공개 모델을 안정적으로 유지하려면 개별 기업이 아니라 산업 전반이 비용과 이해관계를 나누는 공개 모델 컨소시엄이 필요해진다고 주장한다.

Nathan Lambert
#anthropic#nvidia#privacy-design#agent-routing
Training mRNA Language Models Across 25 Species for $165
Article2026년 3월 31일

Training mRNA Language Models Across 25 Species for $165

OpenMed는 단백질 구조 예측, 서열 설계, 코돈 최적화를 잇는 단백질 AI 파이프라인을 구축하고, 코돈 수준 언어모델 비교 끝에 CodonRoBERTa large v2가 생물학적 지표에서 가장 유용하다는 결론을 제시했다.

huggingface.co
#ai-architecture#agent-routing#workflow-automation#llm
Nvidia’s Open Salvo, OpenAI’s Amazon Deal, Grok Cuts Video Prices, and more...
Article2026년 3월 27일

Nvidia’s Open Salvo, OpenAI’s Amazon Deal, Grok Cuts Video Prices, and more...

원문은 AI에 대한 과장된 공포가 규제로 이어질 위험을 경계하며, 동시에 엔비디아가 빠르고 개방적인 에이전트용 오픈 웨이트 모델 Nemotron 3 Super를 공개한 의미를 설명합니다.

deeplearning.ai
#kimi-k2-5#anthropic#nvidia#openai
Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries
Article2026년 3월 10일

Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries

대규모 강화학습의 생성 병목을 해소하려면 추론과 학습을 별도 GPU 풀로 분리하고, 롤아웃 버퍼와 비동기 가중치 동기화를 통해 두 작업이 끊기지 않고 병행되도록 설계해야 한다.

huggingface.co
#ai-architecture#agent-routing#capex-cycle#context-compression
Tricks from OpenAI gpt-oss YOU 🫵 can use with transformers
Article2026년 3월 5일

Tricks from OpenAI gpt-oss YOU 🫵 can use with transformers

OpenAI의 GPT OSS 모델을 transformers에서 효율적으로 실행·미세조정하기 위해 도입된 다운로드형 커널, MXFP4 양자화, Flash Attention 3, 텐서 병렬화 등 핵심 업그레이드를 설명한 글입니다.

huggingface.co
#openai#ai-architecture#agent-memory#agent-routing
이전12342 / 4다음