Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#multimodal
Tag193건YouTube 3Article 190

#multimodal

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#vision-language-models공동문서 179 · 연관도 95%#llm공동문서 173 · 연관도 34%#ai-architecture공동문서 81 · 연관도 27%#semiconductors공동문서 115 · 연관도 24%#agent-routing공동문서 76 · 연관도 23%#agent-memory공동문서 66 · 연관도 20%#nvidia공동문서 47 · 연관도 17%#service-design공동문서 40 · 연관도 15%#workflow-automation공동문서 25 · 연관도 15%#capex-cycle공동문서 32 · 연관도 14%
Exploring Quantization Backends in Diffusers
Article2025년 6월 27일

Exploring Quantization Backends in Diffusers

Diffusers는 Flux의 핵심 구성 요소를 여러 백엔드로 양자화해 이미지 품질을 크게 훼손하지 않으면서 메모리 사용량을 줄일 수 있으며, 정밀도와 백엔드에 따라 속도·메모리·사용 편의성의 차이가 뚜렷하다.

huggingface.co
#ai-architecture#multimodal#agent-memory#context-compression
Transformers backend integration in SGLang
Article2025년 6월 23일

Transformers backend integration in SGLang

SGLang은 트랜스포머스 호환 모델을 자동 대체 백엔드로 실행해 폭넓은 모델 접근성과 고성능 추론·배포 기능을 결합합니다.

huggingface.co
#multimodal#agent-memory#context-compression#retrieval-index
Holo1: New family of GUI automation VLMs powering GUI agent Surfer-H
Article2025년 6월 9일

Holo1: New family of GUI automation VLMs powering GUI agent Surfer-H

H Company는 웹 UI를 이해하고 클릭 위치를 정밀하게 찾는 오픈소스 액션 비전 언어 모델 Holo1과 1,639개 UI 과제로 구성된 WebClick 벤치마크를 공개했으며, 이를 기반으로 브라우저 자동화 에이전트 Surfer H를 구동한다고 밝혔다.

huggingface.co
#ai-architecture#multimodal#workflow-automation#llm
KV Cache from scratch in nanoVLM
Article2025년 6월 4일

KV Cache from scratch in nanoVLM

나노브이엘엠에 계층별 키·값 캐시와 사전 채움·순차 해독 구조를 직접 구현해, 자기회귀 생성의 중복 계산을 줄이고 생성 속도를 38% 높인 과정과 원리를 설명한다.

huggingface.co
#ai-architecture#multimodal#agent-memory#retrieval-index
Introducing AutoRound: Intel’s Advanced Quantization for LLMs and VLMs
Article2025년 5월 27일

Introducing AutoRound: Intel’s Advanced Quantization for LLMs and VLMs

AutoRound는 가중치 반올림과 클리핑 범위를 함께 최적화해 낮은 비트에서도 정확도를 유지하면서 빠른 양자화와 폭넓은 모델·장치·출력 형식 호환성을 제공하는 인텔의 가중치 전용 학습 후 양자화 도구다.

huggingface.co
#multimodal#agent-deployment#ai-infrastructure#capex-cycle
Welcome Llama 4 Maverick & Scout on Hugging Face
Article2025년 5월 22일

Welcome Llama 4 Maverick & Scout on Hugging Face

Meta의 네이티브 멀티모달 MoE 모델 Llama 4 Maverick과 Scout가 출시 당일부터 Hugging Face Hub, Transformers, TGI, 양자화 및 Xet 저장소를 통해 제공되며, 최대 1천만 토큰 문맥과 강력한 추론·이미지·코딩 성능을 지원한다.

huggingface.co
#ai-architecture#multimodal#agent-deployment#agent-routing
AI powers Expedia’s marketing evolution
Article2025년 5월 14일

AI powers Expedia’s marketing evolution

익스피디아 그룹은 인공지능을 분석·콘텐츠 제작·고객 접점 전반에 적용하면서도, 신뢰와 충성도, 인간의 창의성, 부서 간 협업을 여행 마케팅 혁신의 핵심으로 유지하고 있다.

openai.com
#openai#privacy-design#multimodal#search-advertising
LeRobot Community Datasets: The “ImageNet” of Robotics — When and How?
Article2025년 5월 11일

LeRobot Community Datasets: The “ImageNet” of Robotics — When and How?

로봇의 범용화는 모델 구조만의 문제가 아니라 다양한 환경·과제·기체에서 수집된 고품질 데이터를 함께 학습하는 문제이며, LeRobot 공동체는 개방형 로봇 데이터 생태계를 통해 로봇공학의 ‘이미지넷’을 만들어 가고 있다.

huggingface.co
#ai-architecture#multimodal#lerobot#llm
AMIE gains vision: A research AI agent for multimodal diagnostic dialogue
Article2025년 5월 1일

AMIE gains vision: A research AI agent for multimodal diagnostic dialogue

구글 리서치와 딥마인드는 시각 자료를 요청·해석·추론할 수 있는 다중모달 진단 대화 AI 에이전트 AMIE를 공개하고, 시뮬레이션 진료 평가에서 1차 진료의와 비교한 연구 결과를 제시했다.

Google
#google-deepmind#google-research#primary-care-physicians#multimodal
Falcon-Edge: A series of powerful, universal, fine-tunable 1.58bit language models.
Article2025년 5월 1일

Falcon-Edge: A series of powerful, universal, fine-tunable 1.58bit language models.

Falcon Edge는 하나의 사전학습 과정에서 범용 bfloat16 모델, 삼진 가중치 기반 BitNet 모델, 미세조정용 사전 양자화 모델을 함께 제공하는 1.58비트 언어 모델 시리즈다.

huggingface.co
#ai-architecture#multimodal#llm#vision-language-models
Welcoming Llama Guard 4 on Hugging Face Hub
Article2025년 4월 29일

Welcoming Llama Guard 4 on Hugging Face Hub

메타가 공개한 라마 가드 4는 텍스트와 이미지를 함께 검사해 유해한 입력과 출력을 분류하는 120억 매개변수의 다국어 안전 모델이며, 프롬프트 주입과 탈옥을 탐지하는 라마 프롬프트 가드 2도 함께 제공된다.

huggingface.co
#privacy-design#ai-architecture#multimodal#agent-memory
Introducing our latest image generation model in the API
Article2025년 4월 23일

Introducing our latest image generation model in the API

OpenAI는 ChatGPT에서 인기를 얻은 이미지 생성 모델을 gpt image 1 API로 제공해, 개발자와 기업이 고품질 이미지 생성·편집 기능을 제품에 직접 통합할 수 있게 했다고 발표했다.

openai.com
#multimodal#ai-safety#llm#semiconductors
SmolVLA: Efficient Vision-Language-Action Model trained on Lerobot Community Data
Article2025년 4월 21일

SmolVLA: Efficient Vision-Language-Action Model trained on Lerobot Community Data

SmolVLA는 공개 커뮤니티 로봇 데이터로 학습한 4.5억 파라미터 규모의 오픈소스 Vision Language Action 모델로, 저렴한 하드웨어와 소비자급 장비에서도 학습·추론할 수 있도록 설계됐다.

huggingface.co
#lerobot#smolvla#smolvlm2#hugging-face
ScreenSuite - The most comprehensive evaluation suite for GUI Agents!
Article2025년 4월 18일

ScreenSuite - The most comprehensive evaluation suite for GUI Agents!

ScreenSuite는 GUI 에이전트의 지각, 그라운딩, 단일 행동, 다단계 수행 능력을 비전 전용 조건에서 평가하기 위해 13개 벤치마크를 통합한 종합 평가 스위트다.

huggingface.co
#screensuite#smolagents#holo1-7b#hugging-face
The Future of the Global Open-Source AI Ecosystem: From DeepSeek to AI+
Article2025년 4월 18일

The Future of the Global Open-Source AI Ecosystem: From DeepSeek to AI+

딥시크 R1 이후 중국의 인공지능 생태계는 오픈소스를 개별 모델 공개가 아닌 연구·인프라·산업 적용을 연결하는 기본 설계 원칙으로 받아들이며 대규모 배포와 통합 중심의 자생적 구조로 전환했다.

huggingface.co
#multimodal#agent-deployment#agent-routing#llm
Introducing OpenAI o3 and o4-mini
Article2025년 4월 16일

Introducing OpenAI o3 and o4-mini

오픈AI o3와 o4 mini는 강화학습으로 추론 능력과 도구 선택 능력을 함께 확장해, 복잡한 문제를 분석하고 웹 검색·파일 처리·파이썬·이미지 생성 등 여러 도구를 연계해 해결하는 모델이다.

openai.com
#openai#multimodal#agent-routing#llm
Thinking with images
Article2025년 4월 16일

Thinking with images

OpenAI o3와 o4 mini는 이미지를 단순히 인식하는 데 그치지 않고, 확대·회전·자르기와 다른 도구를 활용해 시각 정보와 텍스트를 함께 추론하는 모델이다.

openai.com
#openai#privacy-design#multimodal#ai-safety
Hugging Face to sell open-source robots thanks to Pollen Robotics acquisition 🤖
Article2025년 4월 14일

Hugging Face to sell open-source robots thanks to Pollen Robotics acquisition 🤖

허깅페이스는 오픈소스 휴머노이드 로봇 기업 폴렌 로보틱스를 인수하고, 르로봇 생태계와 리치 2를 결합해 개방형 로봇 소프트웨어·하드웨어를 직접 제공하기 시작했다.

huggingface.co
#multimodal#agent-routing#llm#semiconductors
Introducing GPT-4.1 in the API
Article2025년 4월 14일

Introducing GPT-4.1 in the API

오픈AI는 코딩, 지시 이행, 장문 맥락 처리 능력을 크게 개선하면서 비용과 지연 시간도 낮춘 API 전용 모델군 GPT‑4.1, GPT‑4.1 미니, GPT‑4.1 나노를 공개했다.

openai.com
#openai#privacy-design#multimodal#agent-routing
Get your VLM running in 3 simple steps on Intel CPUs
Article2025년 4월 8일

Get your VLM running in 3 simple steps on Intel CPUs

소형 비전 언어 모델 SmolVLM2를 Optimum Intel과 OpenVINO로 변환·양자화·추론하면 별도 GPU 없이도 Intel CPU에서 지연시간과 처리량을 크게 개선할 수 있다.

huggingface.co
#privacy-design#service-design#multimodal#agent-memory
Smol2Operator: Post-Training GUI Agents for Computer Use
Article2025년 4월 8일

Smol2Operator: Post-Training GUI Agents for Computer Use

Smol2Operator는 작은 비전 언어 모델에 GUI grounding과 행동 추론 능력을 단계적으로 학습시켜, 화면을 이해하고 클릭·입력·스크롤 같은 GUI 행동을 수행하는 에이전트로 발전시키는 공개 재현 가능한 학습 레시피입니다.

huggingface.co
#aguvis#smol2operator#hugging-face#xlangai-aguvis-stage1
SmolVLM2: Bringing Video Understanding to Every Device
Article2025년 4월 8일

SmolVLM2: Bringing Video Understanding to Every Device

SmolVLM2는 2.2B·500M·256M 세 가지 크기로 영상 이해의 높은 메모리 효율과 온디바이스 실행 가능성을 제시하며, Transformers와 MLX를 통해 출시 직후부터 다양한 환경에서 활용할 수 있도록 공개된 소형 비전·영상 언어 모델군이다.

huggingface.co
#multimodal#capex-cycle#context-compression#prompt-library
🚀 Accelerating LLM Inference with TGI on Intel Gaudi
Article2025년 3월 29일

🚀 Accelerating LLM Inference with TGI on Intel Gaudi

Hugging Face는 Intel Gaudi 하드웨어 지원을 Text Generation Inference(TGI) 본체에 네이티브로 통합해, 별도 포크 없이 Gaudi 기반 LLM 추론 배포를 사용할 수 있게 했습니다.

huggingface.co
#ai-architecture#multimodal#llm#semiconductors
Introducing 4o Image Generation
Article2025년 3월 25일

Introducing 4o Image Generation

OpenAI는 GPT 4o에 이미지 생성을 기본 능력으로 통합해 정확한 문자 표현, 세밀한 지시 이행, 대화 기반 수정, 맥락 유지와 사실적인 표현을 갖춘 실용적 시각 커뮤니케이션 도구를 공개했다.

openai.com
#openai#luxury-hospitality#luxury-travel#privacy-design
이전1…567897 / 9다음