Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#vision-language-models
Tag181건YouTube 4Article 177

#vision-language-models

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#multimodal공동문서 179 · 연관도 95%#llm공동문서 165 · 연관도 34%#ai-architecture공동문서 78 · 연관도 27%#agent-routing공동문서 73 · 연관도 23%#semiconductors공동문서 105 · 연관도 22%#agent-memory공동문서 64 · 연관도 20%#nvidia공동문서 45 · 연관도 16%#workflow-automation공동문서 24 · 연관도 15%#capex-cycle공동문서 31 · 연관도 14%#service-design공동문서 36 · 연관도 14%
NVIDIA brings agents to life with DGX Spark and Reachy Mini
Article2025년 12월 15일

NVIDIA brings agents to life with DGX Spark and Reachy Mini

NVIDIA는 CES 2026에서 DGX Spark와 Reachy Mini를 활용해 오픈 모델, 에이전트 프레임워크, 음성·비전·로봇 제어를 결합한 개인형 물리 에이전트 구현 과정을 소개했다.

huggingface.co
#nvidia#dgx-spark#reachy-mini#nemo-agent-toolkit
From Waveforms to Wisdom: The New Benchmark for Auditory Intelligence
Article2025년 12월 3일

From Waveforms to Wisdom: The New Benchmark for Auditory Intelligence

Google Research는 기계 청각 지능을 여덟 가지 핵심 능력으로 표준 평가하는 오픈소스 벤치마크 MSEB를 공개하며, 현재 사운드 임베딩 모델들이 범용성과 견고성에서 큰 성능 여지를 남기고 있다고 밝혔다.

research.google
#multimodal#llm#semiconductors#vision-language-models
Funding grants for new research into AI and mental health
Article2025년 12월 1일

Funding grants for new research into AI and mental health

오픈AI는 인공지능과 정신건강의 위험·편익을 연구하고 안전성과 웰빙을 높이기 위해 독립 연구자들에게 총 200만 달러 규모의 연구비를 지원했으며, 1,000건이 넘는 신청서를 검토해 지원 대상 선정을 마쳤다.

openai.com
#openai#privacy-design#multimodal#llm
Transformers v5: Simple model definitions powering the AI ecosystem
Article2025년 12월 1일

Transformers v5: Simple model definitions powering the AI ecosystem

Transformers v5는 폭발적으로 커진 AI 생태계의 기준 모델 정의 라이브러리로 남기 위해 단순성, 훈련, 추론, 프로덕션 연동, 양자화를 중심으로 구조를 재정비한 릴리스다.

huggingface.co
#ai-architecture#multimodal#agent-deployment#agent-routing
OVHcloud on Hugging Face Inference Providers 🔥
Article2025년 11월 24일

OVHcloud on Hugging Face Inference Providers 🔥

OVHcloud가 Hugging Face Inference Provider로 추가되어 사용자는 Hub와 Python·JavaScript SDK에서 다양한 오픈 웨이트 모델을 OVHcloud의 서버리스 추론 환경으로 호출할 수 있게 됐다.

huggingface.co
#service-design#multimodal#agent-routing#llm
Introducing AnyLanguageModel: One API for Local and Remote LLMs on Apple Platforms
Article2025년 11월 20일

Introducing AnyLanguageModel: One API for Local and Remote LLMs on Apple Platforms

AnyLanguageModel은 Apple 플랫폼에서 로컬 모델과 클라우드 LLM을 하나의 Swift API로 다루게 해 통합 부담을 줄이려는 패키지다.

huggingface.co
#anthropic#apple-silicon#privacy-design#multimodal
Generative UI: A rich, custom, visual interactive user experience for any prompt
Article2025년 11월 18일

Generative UI: A rich, custom, visual interactive user experience for any prompt

구글은 사용자의 어떤 프롬프트에도 맞춰 웹페이지, 도구, 시뮬레이션 같은 맞춤형 인터랙티브 경험을 즉석 생성하는 생성형 UI 구현을 소개하고, Gemini 앱과 Google Search AI Mode에 실험적으로 적용한다고 밝혔다.

research.google
#multimodal#context-compression#prompt-library#search-advertising
Accelerating the magic cycle of research breakthroughs and real-world applications
Article2025년 10월 31일

Accelerating the magic cycle of research breakthroughs and real-world applications

Google Research는 지구 관측, 암 유전체, 양자 컴퓨팅, 의료·과학 AI 등에서 연구 성과와 실제 적용이 서로를 밀어 올리는 ‘연구의 마법적 순환’이 빠르게 가속되고 있다고 설명한다.

research.google
#privacy-design#multimodal#ai-distribution#search-advertising
Google Earth AI: Unlocking geospatial insights with foundation models and cross-modal reasoning
Article2025년 10월 23일

Google Earth AI: Unlocking geospatial insights with foundation models and cross-modal reasoning

Google Earth AI는 위성영상, 인구·이동, 환경 예측 모델과 Gemini 기반 지리공간 추론 에이전트를 결합해 복잡한 현실 문제를 단계적으로 분석하고 실행 가능한 지리공간 인사이트를 제공하려는 Google의 지리공간 AI 체계다.

research.google
#gemini#google-research#google-earth-ai#remote-sensing-foundations
Sentence Transformers is joining Hugging Face!
Article2025년 10월 22일

Sentence Transformers is joining Hugging Face!

문장 임베딩 생태계의 핵심 오픈소스 프로젝트인 센텐스 트랜스포머스가 다름슈타트 공과대학교의 유비쿼터스 지식 처리 연구실에서 허깅페이스로 이관되며, 기존 개방성과 공동체 중심 운영을 유지한 채 새로운 성장 단계에 들어선다.

huggingface.co
#ai-architecture#multimodal#llm#semiconductors
Teaching Gemini to spot exploding stars with just a few examples
Article2025년 10월 20일

Teaching Gemini to spot exploding stars with just a few examples

제미나이는 각 천문 관측 조사별 15개의 주석 예시만으로 초신성 같은 일시적 천체 현상을 93% 정확도로 분류하고, 판단 근거와 추적 관측 우선순위를 자연어로 설명하는 천문 보조 도구로 활용될 수 있음을 보였다.

research.google
#multimodal#agent-routing#context-compression#prompt-library
Unlock the power of images with AI Sheets
Article2025년 10월 16일

Unlock the power of images with AI Sheets

허깅페이스 AI Sheets는 스프레드시트 안에서 오픈 비전 모델을 활용해 이미지의 정보를 추출·구조화하고, 텍스트와 이미지를 생성·편집하며, 완성된 데이터셋까지 내보낼 수 있게 확장됐다.

huggingface.co
#service-design#multimodal#agent-routing#prompt-library
AI for Food Allergies
Article2025년 10월 15일

AI for Food Allergies

식품 알레르기 연구에 인공지능을 접목하고, 분자·임상·식품 안전 데이터를 개방형으로 연결해 예측과 진단, 치료제 탐색, 소비자 보호를 발전시키려는 공동체 주도 연구 프로젝트를 소개한다.

huggingface.co
#service-design#multimodal#llm#vision-language-models
Coral NPU: A full-stack platform for Edge AI
Article2025년 10월 15일

Coral NPU: A full-stack platform for Edge AI

구글 리서치는 Coral NPU를 저전력 엣지·웨어러블 기기에서 항상 켜진 개인 AI를 구현하기 위한 개방형 풀스택 NPU 플랫폼으로 소개했다.

research.google
#privacy-design#ai-architecture#multimodal#search-advertising
We Got Claude to Fine-Tune an Open Source LLM
Article2025년 10월 14일

We Got Claude to Fine-Tune an Open Source LLM

이 글은 Hugging Face Skills의 hf llm trainer를 통해 Claude 같은 코딩 에이전트가 데이터 검증, GPU 선택, 학습 작업 제출, 모니터링, Hub 업로드까지 오픈소스 LLM 파인튜닝 전 과정을 처리하는 방법을 설명한다.

huggingface.co
#privacy-design#multimodal#agent-deployment#agent-routing
A collaborative approach to image generation
Article2025년 10월 2일

A collaborative approach to image generation

Google Research는 사용자의 선택을 여러 턴에 걸쳐 학습해 텍스트 이미지 결과를 점진적으로 개선하는 강화학습 에이전트 PASTA를 소개했다.

research.google
#privacy-design#multimodal#context-compression#prompt-library
Introducing RTEB: A New Standard for Retrieval Evaluation
Article2025년 10월 1일

Introducing RTEB: A New Standard for Retrieval Evaluation

RTEB는 공개·비공개 데이터셋을 함께 활용해 임베딩 모델의 미지 데이터 검색 성능과 실제 응용 적합성을 공정하게 평가하려는 새로운 검색 벤치마크다.

huggingface.co
#multimodal#agent-memory#change-management#organizational-redesign
The anatomy of a personal health agent
Article2025년 9월 30일

The anatomy of a personal health agent

Google Research는 웨어러블 데이터와 건강 데이터를 결합해 개인화된 근거 기반 건강 인사이트와 코칭을 제공하는 다중 에이전트 연구 프레임워크 PHA를 제안하고, 데이터 과학·도메인 전문가·헬스 코치 역할을 분리해 평가했다.

research.google
#service-design#ai-architecture#multimodal#agent-routing
AfriMed-QA: Benchmarking large language models for global health
Article2025년 9월 24일

AfriMed-QA: Benchmarking large language models for global health

AfriMed QA는 아프리카 의료 맥락에 맞춘 대규모 질의응답 벤치마크로, LLM이 지역별 질병·언어·문화·의료 환경 차이를 얼마나 잘 처리하는지 평가하기 위해 구축됐다.

research.google
#multimodal#llm#semiconductors#vision-language-models
Scaleway on Hugging Face Inference Providers 🔥
Article2025년 9월 19일

Scaleway on Hugging Face Inference Providers 🔥

Scaleway가 Hugging Face 추론 제공자로 추가되면서 사용자는 허브와 파이썬·자바스크립트 SDK에서 다양한 공개 가중치 모델을 Scaleway의 서버리스 추론 환경으로 이용할 수 있게 되었습니다.

huggingface.co
#service-design#multimodal#agent-routing#llm
How to Build a Healthcare Robot from Simulation to Deployment with NVIDIA Isaac for Healthcare
Article2025년 9월 17일

How to Build a Healthcare Robot from Simulation to Deployment with NVIDIA Isaac for Healthcare

이 글은 NVIDIA Isaac for Healthcare의 SO ARM 스타터 워크플로를 통해 시뮬레이션 데이터 수집, GR00T N1.5 후학습, Isaac Lab 평가, 실제 SO ARM101 하드웨어 배포까지 이어지는 의료 로봇 개발 과정을 설명한다.

huggingface.co
#nvidia#ai-architecture#multimodal#agent-deployment
Post-Training Isaac GR00T N1.5 for LeRobot SO-101 Arm
Article2025년 9월 17일

Post-Training Isaac GR00T N1.5 for LeRobot SO-101 Arm

이 글은 NVIDIA Isaac GR00T N1.5를 저가형 오픈소스 LeRobot SO 101 로봇팔에 맞게 후학습하고, 데이터 준비부터 평가·실물 배포까지 진행하는 절차를 단계별로 설명한다.

huggingface.co
#service-design#multimodal#agent-deployment#prompt-library
Learn Your Way: Reimagining textbooks with generative AI
Article2025년 9월 16일

Learn Your Way: Reimagining textbooks with generative AI

구글 리서치의 Learn Your Way는 생성형 AI로 교과서를 개인화된 다중 형식 학습 경험으로 바꾸고, 초기 연구에서 디지털 리더보다 학습 성과와 유지 점수를 높인 교육 실험이다.

research.google
#multimodal#agent-routing#llm#semiconductors
SafetyKit scales risk agents with OpenAI’s most capable models
Article2025년 9월 9일

SafetyKit scales risk agents with OpenAI’s most capable models

SafetyKit은 OpenAI의 GPT 5, GPT 4.1, deep research, CUA를 조합해 사기·규정 위반·위험 콘텐츠를 멀티모달로 검토하는 전용 에이전트를 확장하고, 고객 콘텐츠 100% 검토에서 95% 이상 정확도를 보고했다.

openai.com
#openai#safetykit#gpt-5#gpt-4-1
이전1…3456785 / 8다음