Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#vision-language-models
Tag181건YouTube 4Article 177

#vision-language-models

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#multimodal공동문서 179 · 연관도 95%#llm공동문서 165 · 연관도 34%#ai-architecture공동문서 78 · 연관도 27%#agent-routing공동문서 73 · 연관도 23%#semiconductors공동문서 105 · 연관도 22%#agent-memory공동문서 64 · 연관도 20%#nvidia공동문서 45 · 연관도 16%#workflow-automation공동문서 24 · 연관도 15%#capex-cycle공동문서 31 · 연관도 14%#service-design공동문서 36 · 연관도 14%
Anthropic vs. the U.S. Government, Nano Banana’s Makeover, Frontier Agent Management, and more...
Article2026년 3월 6일

Anthropic vs. the U.S. Government, Nano Banana’s Makeover, Frontier Agent Management, and more...

본문은 코딩 에이전트에 최신 API 문맥을 제공하는 Context Hub 발표, Google Nano Banana 2의 가격·속도 개선, 그리고 Anthropic·OpenAI·미국 전쟁부 사이의 군사용 AI 계약 갈등을 다룬다.

deeplearning.ai
#anthropic#service-design#ai-architecture#multimodal
How Balyasny Asset Management built an AI research engine
Article2026년 3월 6일

How Balyasny Asset Management built an AI research engine

발리아스니 자산운용은 엄격한 모델 평가, 현업 중심의 피드백, 중앙 플랫폼과 팀별 맞춤형 에이전트를 결합해 며칠 걸리던 투자 리서치를 수 시간 안에 수행하는 인공지능 연구 체계를 구축했다.

openai.com
#ai-architecture#multimodal#agent-deployment#llm
Gemma 3n fully available in the open-source ecosystem!
Article2026년 3월 6일

Gemma 3n fully available in the open-source ecosystem!

젬마 3n은 적은 GPU 메모리로 구동되는 네이티브 멀티모달 모델로, 트랜스포머스·MLX·라마닷씨피피·트랜스포머스닷제이에스 등 주요 오픈소스 생태계에서 추론과 미세조정을 지원한다.

huggingface.co
#ai-architecture#multimodal#agent-memory#agent-routing
NVIDIA Releases 6 Million Multi-Lingual Reasoning Dataset
Article2026년 3월 5일

NVIDIA Releases 6 Million Multi-Lingual Reasoning Dataset

NVIDIA는 기존 영어 추론 데이터를 프랑스어·독일어·이탈리아어·일본어·스페인어로 확장하고, 번역 환각과 언어 이탈을 줄이기 위한 필터링 절차를 적용한 600만 건 규모의 Nemotron 다국어 추론 데이터셋을 공개했다.

huggingface.co
#nvidia#service-design#ai-architecture#multimodal
뇌과학자가 말하는 AI 시대에도 끝까지 살아남는 5가지 능력 ㅣ 김대식 교수님의 고전 책 5권 추천 ㅣ도서리뷰
YouTube2026년 3월 4일

뇌과학자가 말하는 AI 시대에도 끝까지 살아남는 5가지 능력 ㅣ 김대식 교수님의 고전 책 5권 추천 ㅣ도서리뷰

인간의 끝까지 남는 경쟁력은 더 나은 계산이 아니라, 불확실성 속에서 방향을 고르고 손실 뒤에 다시 귀환하며 타인을 서사 있는 주체로 대하는 유연성·회복탄력성·목표 설정·연민의 판단 구조다.

책과삶
#openclaw#llm#ai-architecture#vision-language-models
NVIDIA Cosmos Reason 2 Brings Advanced Reasoning To Physical AI
Article2026년 3월 3일

NVIDIA Cosmos Reason 2 Brings Advanced Reasoning To Physical AI

NVIDIA는 물리 AI를 위한 공개 추론 비전 언어 모델 Cosmos Reason 2를 공개하며, 시공간 이해·긴 문맥·로봇 계획·영상 분석 역량을 크게 강화했다고 밝혔다.

huggingface.co
#nvidia#multimodal#ai-infrastructure#capex-cycle
🎧 How OpenAI’s Codex Team Uses Their Coding Agent
Article2026년 2월 18일

🎧 How OpenAI’s Codex Team Uses Their Coding Agent

OpenAI Codex 팀은 전용 GUI, 자동화와 스킬, 빠른 모델, 리뷰 보조 기능을 통해 코딩 에이전트를 단순 코드 생성 도구가 아니라 개발 흐름 전체를 바꾸는 작업 환경으로 만들고 있다고 설명한다.

Rhea Purohit
#openai#privacy-design#multimodal#agent-routing
Transcript: 'How OpenAI’s Codex Team Uses Their Coding Agent
Article2026년 2월 18일

Transcript: 'How OpenAI’s Codex Team Uses Their Coding Agent

OpenAI Codex 팀은 Codex 앱과 O3 Codex를 통해 전문 개발 생산성을 밀어붙이는 동시에, 멀티태스킹과 장기 실행이 가능한 에이전트 경험을 더 넓은 기술 사용자층에 열고 있다고 설명한다.

Dan Shipper
#anthropic#openai#privacy-design#multimodal
Diffusers welcomes FLUX-2
Article2026년 2월 17일

Diffusers welcomes FLUX-2

FLUX.2는 처음부터 새로 사전학습한 이미지 생성·편집 모델로, Diffusers는 새로운 구조와 다중 이미지 참조 기능부터 8~80GB급 그래픽 메모리 환경별 추론 방법까지 구체적으로 소개한다.

huggingface.co
#ai-architecture#multimodal#agent-memory#agent-routing
Teaching AI to read a map
Article2026년 2월 17일

Teaching AI to read a map

구글 연구진은 합성 지도와 경로 주석을 대규모로 생성하는 MapTrace 파이프라인을 제안해, 멀티모달 언어모델이 지도 위에서 벽과 통로의 제약을 지키며 경로를 추적하는 공간 추론 능력을 학습할 수 있음을 보였다.

research.google
#multimodal#maptrace#agent-routing#prompt-library
Introducing GPT-5.3-Codex-Spark
Article2026년 2월 12일

Introducing GPT-5.3-Codex-Spark

GPT 5.3 Codex Spark는 초당 1,000개 이상의 토큰을 생성하는 초저지연 하드웨어를 기반으로, 코딩 과정에서 즉각적인 수정과 반복 작업을 지원하도록 설계된 실시간 코딩 모델의 연구용 미리보기다.

openai.com
#multimodal#agent-routing#workflow-automation#llm
Beyond one-on-one: Authoring, simulating, and testing dynamic human-AI group conversations
Article2026년 2월 10일

Beyond one-on-one: Authoring, simulating, and testing dynamic human-AI group conversations

DialogLab은 일대일 LLM 대화의 한계를 넘어, 역할·집단 구조·발화 순서·즉흥성을 함께 다루는 인간 AI 다자간 대화 제작·시뮬레이션·검증용 오픈소스 연구 프로토타입이다.

research.google
#multimodal#agent-routing#llm#semiconductors
How AI tools can redefine universal design to increase accessibility
Article2026년 2월 5일

How AI tools can redefine universal design to increase accessibility

Google Research는 장애 커뮤니티와의 공동 설계를 바탕으로, 사용자의 맥락과 필요에 맞춰 스스로 조정되는 멀티모달 AI 기반 ‘Natively Adaptive Interfaces’가 보편적 디자인과 접근성을 새롭게 정의할 수 있다고 설명한다.

research.google
#privacy-design#ai-architecture#multimodal#agent-routing
Collaborating on a nationwide randomized study of AI in real-world virtual care
Article2026년 2월 3일

Collaborating on a nationwide randomized study of AI in real-world virtual care

Google은 Included Health와 함께 실제 가상진료 환경에서 대화형 의료 AI의 안전성·유용성·한계를 평가하기 위한 전국 단위 무작위 연구를 추진한다.

research.google
#ai-architecture#multimodal#agent-routing#llm
Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents
Article2026년 1월 30일

Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents

NVIDIA Nemotron 3 Nano Omni는 문서, 이미지, 비디오, 오디오를 긴 컨텍스트 안에서 함께 이해하도록 설계된 오픈 웨이트 옴니모달 모델로, 문서 지능·영상/음성 이해·GUI 에이전트 작업에서 높은 정확도와 효율을 내세운다.

huggingface.co
#mamba#nvidia#hugging-face#c-radiov4-h
Differential Transformer V2
Article2026년 1월 20일

Differential Transformer V2

DIFF V2는 같은 GQA 그룹의 두 쿼리 헤드 출력을 토큰·헤드별 계수로 차감해, 키·값 헤드와 표준 어텐션 커널을 그대로 유지하면서 디코딩 효율, 학습 안정성, 표현 자유도, 출력 투영의 매개변수 효율을 함께 개선한 구조다.

huggingface.co
#ai-architecture#multimodal#llm#semiconductors
ServiceNow powers actionable enterprise AI with OpenAI
Article2026년 1월 20일

ServiceNow powers actionable enterprise AI with OpenAI

ServiceNow와 OpenAI는 다년 계약을 통해 OpenAI의 프런티어 모델과 멀티모달 기능을 ServiceNow의 기업 워크플로에 직접 결합해, 기업 AI가 답변을 넘어 실제 업무 실행까지 수행하도록 하겠다고 발표했다.

openai.com
#openai#privacy-design#service-design#multimodal
Next generation medical image interpretation with MedGemma 1.5 and medical speech to text with MedASR
Article2026년 1월 13일

Next generation medical image interpretation with MedGemma 1.5 and medical speech to text with MedASR

Google Research는 의료 영상 해석을 강화한 공개 모델 MedGemma 1.5 4B와 의료 음성 인식 모델 MedASR을 발표하며, 개발자가 의료 AI 애플리케이션을 평가·조정·확장할 수 있는 기반을 넓혔다.

research.google
#privacy-design#multimodal#agent-routing#llm
Investing in Performance: Fine-tune small models with LLM insights - a CFM case study
Article2026년 1월 12일

Investing in Performance: Fine-tune small models with LLM insights - a CFM case study

CFM 사례는 금융 뉴스 NER에서 대형 LLM을 직접 쓰기보다 LLM으로 라벨을 만들고 검수한 뒤 소형 모델을 미세조정하면 정확도와 비용 효율을 함께 개선할 수 있음을 보여준다.

huggingface.co
#service-design#ai-architecture#multimodal#agent-deployment
Cohere on Hugging Face Inference Providers 🔥
Article2026년 1월 9일

Cohere on Hugging Face Inference Providers 🔥

코히어가 허깅페이스 허브의 추론 제공자로 합류하면서 기업용 언어·검색·다국어·멀티모달 모델 9종을 웹 화면과 여러 클라이언트 라이브러리에서 서버리스 방식으로 사용할 수 있게 됐다.

huggingface.co
#multimodal#capex-cycle#llm#semiconductors
How Tolan builds voice-first AI with GPT-5.1
Article2026년 1월 7일

How Tolan builds voice-first AI with GPT-5.1

Tolan은 GPT 5.1의 낮은 지연시간과 높은 지시 이행 능력을 바탕으로, 매 턴 재구성되는 문맥과 정교한 기억·캐릭터 시스템을 결합해 자연스럽고 일관된 음성형 AI 동반자를 구현했다.

openai.com
#ai-architecture#multimodal#context-compression#prompt-library
Architectural Choices in China's Open-Source AI Ecosystem: Building Beyond DeepSeek
Article2025년 12월 23일

Architectural Choices in China's Open-Source AI Ecosystem: Building Beyond DeepSeek

중국의 오픈소스 인공지능 생태계는 딥시크 이후 최고 성능의 단일 모델보다 전문가 혼합 구조, 다중양식, 소형 모델, 허용적 라이선스, 자국산 하드웨어와 배포 도구를 결합한 지속 가능한 시스템 구축으로 경쟁의 중심을 옮기고 있다.

huggingface.co
#china#ai-architecture#multimodal#agent-deployment
Deepening our collaboration with the U.S. Department of Energy
Article2025년 12월 18일

Deepening our collaboration with the U.S. Department of Energy

OpenAI와 미국 에너지부(DOE)는 AI와 첨단 컴퓨팅을 활용해 과학 발견을 가속하기 위한 협력 양해각서(MOU)를 체결하고, 국가연구소와의 기존 협력을 더 구조화된 방식으로 확대하기로 했다.

openai.com
#openai#privacy-design#multimodal#agent-routing
The new ChatGPT Images is here
Article2025년 12월 16일

The new ChatGPT Images is here

OpenAI는 더 정밀한 이미지 편집, 향상된 지시 이행, 작은 글자 렌더링, 최대 4배 빠른 생성 속도를 갖춘 새 ChatGPT Images와 API 모델 GPT Image 1.5를 공개했습니다.

openai.com
#openai#privacy-design#multimodal#agent-routing
이전123456…84 / 8다음