Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#vision-language-models
Tag181건YouTube 4Article 177

#vision-language-models

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#multimodal공동문서 179 · 연관도 95%#llm공동문서 165 · 연관도 34%#ai-architecture공동문서 78 · 연관도 27%#agent-routing공동문서 73 · 연관도 23%#semiconductors공동문서 105 · 연관도 22%#agent-memory공동문서 64 · 연관도 20%#nvidia공동문서 45 · 연관도 16%#workflow-automation공동문서 24 · 연관도 15%#capex-cycle공동문서 31 · 연관도 14%#service-design공동문서 36 · 연관도 14%
Boosting multimodal inference performance by >10% with a single Python dictionary
Article2026년 5월 4일

Boosting multimodal inference performance by >10% with a single Python dictionary

Modal은 SGLang의 멀티모달 추론 스케줄러에서 반복적인 CUDA IPC 핸들 열기 비용을 Python dict 캐시로 제거해 Qwen2.5 VL 3B Instruct 단일 H100 벤치마크에서 처리량 16.2%, 평균 지연 10% 이상 개선했다고 설명한다.

Modal
#modal#pytorch#sglang#cuda-ipc
Granite Embedding Multilingual R2: Open Apache 2.0 Multilingual Embeddings with 32K Context — Best Sub-100M Retrieval Quality
Article2026년 5월 4일

Granite Embedding Multilingual R2: Open Apache 2.0 Multilingual Embeddings with 32K Context — Best Sub-100M Retrieval Quality

IBM의 Granite Embedding Multilingual R2는 200개 이상의 언어와 32K 토큰 문맥을 지원하며, 97M 소형 모델과 311M 고성능 모델로 다국어·장문·코드 검색의 성능과 배포 효율을 함께 높인 Apache 2.0 임베딩 모델군이다.

huggingface.co
#ai-architecture#multimodal#agent-memory#retrieval-index
Catalyzing scientific impact through global partnerships and open resources
Article2026년 5월 1일

Catalyzing scientific impact through global partnerships and open resources

Google Research Science 팀은 책임 있고 포용적인 개방 과학 원칙 아래 오픈소스 도구, 공개 데이터셋, 글로벌 파트너십을 통해 유전체학·뇌과학·기후·보건·생물다양성 분야의 실제 연구 성과와 사회적 영향을 확장하고 있다고 설명한다.

research.google
#ai-architecture#multimodal#agent-routing#search-advertising
GPT-5.5 Outperforms (and Hallucinates), Kimi K2.6 Leads Open LLMs, AI Strains Climate Pledges, and more...
Article2026년 5월 1일

GPT-5.5 Outperforms (and Hallucinates), Kimi K2.6 Leads Open LLMs, AI Strains Climate Pledges, and more...

글은 최신 AI 모델 사용법 교육 안내를 시작으로, GPT 5.5의 높은 벤치마크 성과와 환각 문제, 대형 AI 기업의 데이터센터 확장으로 흔들리는 탄소 감축 약속, 그리고 오픈 가중치 모델 Kimi K2.6의 경쟁력을 함께 다룬다.

@DeepLearningAI
#kimi-k2-5#anthropic#agent-swarms#service-design
Granite 4.0 3B Vision: Compact Multimodal Intelligence for Enterprise Documents
Article2026년 4월 30일

Granite 4.0 3B Vision: Compact Multimodal Intelligence for Enterprise Documents

Granite 4.0 3B Vision은 복잡한 기업 문서의 표·차트·핵심 값 쌍을 정밀하게 추출하도록 설계된 30억 매개변수급 소형 비전·언어 모델이다.

huggingface.co
#ai-architecture#multimodal#agent-routing#workflow-automation
Demis Hassabis: Agents, AGI & The Next Big Scientific Breakthrough
YouTube2026년 4월 29일

Demis Hassabis: Agents, AGI & The Next Big Scientific Breakthrough

지속 학습, 장기 추론, 기억의 일부 측면은 아직 해결되지 않았고, 이런 능력들은 AGI에 필수적인 구성 요소로 남아 있다

Y Combinator
#alphafold#deepmind#vision-language-models#change-management
Choco automates food distribution with AI agents
Article2026년 4월 27일

Choco automates food distribution with AI agents

Choco는 OpenAI API를 기반으로 이메일·문자·음성·이미지 등 다양한 주문 입력을 구조화된 ERP 주문으로 자동 변환하며, 글로벌 식품 유통망의 수작업 병목을 줄이고 상시 운영 체계를 구축하고 있다.

openai.com
#choco#orderagent#voiceagent#openai-api
GLM 5.1 Thinks Strategically, Data-Center Revolt Intensifies, When Helpful LLMs Turn Unhelpful, and more...
Article2026년 4월 24일

GLM 5.1 Thinks Strategically, Data-Center Revolt Intensifies, When Helpful LLMs Turn Unhelpful, and more...

본문은 코딩 에이전트가 프론트엔드에는 큰 가속을 주지만 백엔드·인프라·연구로 갈수록 한계가 커진다는 판단과, 장시간 자율 작업을 지향하는 GLM 5.1 및 초기 산업 현장의 휴머노이드 로봇 배치를 다룬다.

@DeepLearningAI
#anthropic#ai-architecture#multimodal#agent-routing
The PR you would have opened yourself
Article2026년 4월 16일

The PR you would have opened yourself

transformers 모델을 mlx lm으로 빠르고 정확하게 이식하도록 돕는 스킬과 비에이전트 테스트 하네스를 구축하되, 코드 소유권과 최종 판단은 기여자와 리뷰어에게 남겨 두는 접근을 설명한다.

huggingface.co
#privacy-design#ai-architecture#multimodal#llm
Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers
Article2026년 4월 16일

Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers

이 글은 Sentence Transformers로 텍스트와 이미지 등 여러 모달리티를 다루는 임베딩 모델을 자체 데이터에 맞게 파인튜닝하는 방법을, 시각 문서 검색 사례를 중심으로 설명합니다.

huggingface.co
#ai-architecture#multimodal#llm#vision-language-models
Introducing Fire-PDF: Firecrawl's New PDF Parsing Engine
Article2026년 4월 14일

Introducing Fire-PDF: Firecrawl's New PDF Parsing Engine

Firecrawl은 PDF를 구조화된 마크다운으로 더 빠르고 정확하게 변환하기 위해 Rust 기반 새 PDF 파싱 엔진 Fire PDF를 출시했으며, 모든 API PDF 처리에 자동 적용한다고 밝혔다.

Eric Ciarla
#multimodal#ai-infrastructure#capex-cycle#llm
State of Open Source on Hugging Face: Spring 2026
Article2026년 4월 14일

State of Open Source on Hugging Face: Spring 2026

허깅페이스의 오픈소스 인공지능 생태계는 사용자·모델·데이터셋과 파생 창작물이 급증한 가운데, 중국과 독립 개발자의 영향력 확대, 소형 모델 중심의 실용적 채택, 국가 주권과 지역 생태계의 부상을 동시에 보여준다.

huggingface.co
#ai-architecture#multimodal#agent-deployment#agent-routing
Open ASR Leaderboard: Trends and Insights with New Multilingual & Long-Form Tracks
Article2026년 4월 13일

Open ASR Leaderboard: Trends and Insights with New Multilingual & Long-Form Tracks

개방형 자동 음성 인식 순위표는 정확도와 처리 속도뿐 아니라 다국어·장문 전사 성능까지 비교하며, 언어 범용성·특화 정확도·처리량 사이의 뚜렷한 상충 관계를 보여준다.

huggingface.co
#multimodal#agent-routing#llm#semiconductors
Anthropic’s Claude Mythos Problem, Dark DNA Unveiled, Pitfalls for Assistive Models, and more...
Article2026년 4월 10일

Anthropic’s Claude Mythos Problem, Dark DNA Unveiled, Pitfalls for Assistive Models, and more...

이 글은 AI 코딩 에이전트가 소프트웨어 엔지니어링과 고용 논의를 바꾸는 흐름을 짚고, Anthropic의 Claude Mythos Preview가 제기한 사이버보안 위험과 시각장애인을 위한 보조 AI의 심리적·사회적 함의를 함께 다룬다.

deeplearning.ai
#anthropic#ai-architecture#multimodal#agent-routing
Multimodal Embedding & Reranker Models with Sentence Transformers
Article2026년 4월 10일

Multimodal Embedding & Reranker Models with Sentence Transformers

Sentence Transformers 5.4는 텍스트·이미지·오디오·비디오를 같은 인터페이스로 임베딩하고, 서로 다른 형식의 입력 쌍을 재평가해 교차 모달 검색과 멀티모달 검색 증강 생성 파이프라인을 구축할 수 있게 한다.

huggingface.co
#service-design#ai-architecture#multimodal#llm
Claude Code’s Source Leaks, OpenAI Exits Video Generation, Gemini Adds Music Generation, and more...
Article2026년 4월 3일

Claude Code’s Source Leaks, OpenAI Exits Video Generation, Gemini Adds Music Generation, and more...

원문은 음성 UI가 AI 애플리케이션의 중요한 인터페이스가 될 것이라는 전망을 중심으로, Claude Code 소스 유출이 드러낸 에이전트 구조와 OpenAI의 Sora 종료 계획을 함께 다룬다.

@DeepLearningAI
#anthropic#openai#claude-code#linear-attention
Welcome Gemma 4: Frontier multimodal intelligence on device
Article2026년 4월 2일

Welcome Gemma 4: Frontier multimodal intelligence on device

Gemma 4는 이미지·텍스트·오디오를 처리하면서 긴 문맥, 추론 효율, 기기 내 배포를 함께 겨냥한 Apache 2.0 기반 개방형 멀티모달 모델군이다.

huggingface.co
#service-design#token-efficiency#multimodal#travel-hospitality
Mapping the modern world: How S2Vec learns the language of our cities
Article2026년 3월 24일

Mapping the modern world: How S2Vec learns the language of our cities

구글 리서치의 S2Vec은 도로·건물·상점 같은 구축환경 데이터를 S2 셀 기반 이미지처럼 변환하고 자기지도 학습으로 임베딩해 전 세계 사회경제·환경 패턴 예측에 활용하는 지리 AI 프레임워크다.

research.google
#privacy-design#multimodal#llm#semiconductors
Google Research at The Check Up: from healthcare innovation to real-world care settings
Article2026년 3월 17일

Google Research at The Check Up: from healthcare innovation to real-world care settings

구글 리서치는 The Check Up에서 개인 맞춤형 건강 관리, 임상의 협업, 개발자 생태계, 공중보건, 생명과학 연구 전반에 AI를 책임 있게 적용해 실제 의료 현장으로 옮기려는 최신 연구 성과를 소개했다.

research.google
#service-design#multimodal#search-advertising#zero-click-search
Granite 4.1 LLMs: How They’re Built
Article2026년 3월 17일

Granite 4.1 LLMs: How They’re Built

Granite 4.1은 3B·8B·30B 밀집형 디코더 전용 LLM 제품군으로, 약 15조 토큰의 5단계 사전학습, 512K 장문 컨텍스트 확장, 약 410만 개 고품질 SFT 샘플, 다단계 강화학습을 통해 수학·코딩·지시수행·대화 성능을 강화한 모델이다.

huggingface.co
#ai-architecture#multimodal#agent-routing#prompt-library
Holotron-12B - High Throughput Computer Use Agent
Article2026년 3월 17일

Holotron-12B - High Throughput Computer Use Agent

홀로트론 12B는 하이브리드 상태공간모델과 어텐션 구조를 기반으로 긴 문맥과 다중 이미지를 효율적으로 처리하면서 높은 컴퓨터 사용 성능과 추론 처리량을 달성한 120억 매개변수 멀티모달 에이전트 모델이다.

huggingface.co
#nvidia#ai-architecture#multimodal#agent-memory
Exploring the feasibility of conversational diagnostic AI in a real-world clinical study
Article2026년 3월 11일

Exploring the feasibility of conversational diagnostic AI in a real-world clinical study

구글 리서치·구글 딥마인드와 Beth Israel Deaconess Medical Center의 전향적 단일기관 연구는 대화형 의료 AI AMIE가 실제 일차진료 방문 전 병력 청취를 감독하에 수행하는 것이 초기 단계에서 실행 가능하고 안전하게 수용될 수 있음을 보고했다.

research.google
#multimodal#agent-routing#semiconductors#vision-language-models
Wayfair boosts catalog accuracy and support speed with OpenAI
Article2026년 3월 11일

Wayfair boosts catalog accuracy and support speed with OpenAI

Wayfair는 OpenAI 모델을 상품 카탈로그와 공급업체 지원 업무에 내장해 수백만 개 상품의 속성 정확도를 높이고 티켓 처리 속도를 개선했다.

openai.com
#openai#privacy-design#ai-architecture#multimodal
LeRobot v0.5.0: Scaling Every Dimension
Article2026년 3월 9일

LeRobot v0.5.0: Scaling Every Dimension

LeRobot v0.5.0은 휴머노이드와 다양한 로봇 하드웨어, 새로운 VLA 정책, 고속 데이터 처리, Hub 기반 시뮬레이션 환경, 현대화된 개발 기반을 한꺼번에 확장한 역대 최대 규모의 릴리스다.

huggingface.co
#nvidia#multimodal#agent-routing#ai-infrastructure
이전12345…83 / 8다음