Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#agent-memory
Tag543건YouTube 19Article 524

#agent-memory

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#retrieval-index공동문서 248 · 연관도 66%#context-compression공동문서 286 · 연관도 60%#semiconductors공동문서 457 · 연관도 56%#applications공동문서 432 · 연관도 56%#llm공동문서 468 · 연관도 55%#agent-routing공동문서 201 · 연관도 37%#ai-architecture공동문서 152 · 연관도 30%#agent-deployment공동문서 98 · 연관도 23%#vision-language-models공동문서 64 · 연관도 20%#multimodal공동문서 66 · 연관도 20%
KV Cache from scratch in nanoVLM
Article2025년 6월 4일

KV Cache from scratch in nanoVLM

나노브이엘엠에 계층별 키·값 캐시와 사전 채움·순차 해독 구조를 직접 구현해, 자기회귀 생성의 중복 계산을 줄이고 생성 속도를 38% 높인 과정과 원리를 설명한다.

huggingface.co
#ai-architecture#multimodal#agent-memory#retrieval-index
(LoRA) Fine-Tuning FLUX.1-dev on Consumer Hardware
Article2025년 6월 1일

(LoRA) Fine-Tuning FLUX.1-dev on Consumer Hardware

FLUX.1 dev의 트랜스포머를 4비트로 양자화하고 LoRA, 8비트 옵티마이저, 체크포인팅, 잠재 표현·텍스트 임베딩 캐시를 결합해 단일 GPU에서 약 10GB 미만의 VRAM으로 미세 조정하는 방법을 설명한다.

huggingface.co
#ai-architecture#agent-memory#capex-cycle#context-compression
CodeAgents + Structure: A Better Way to Execute Actions
Article2025년 5월 28일

CodeAgents + Structure: A Better Way to Execute Actions

이 글은 CodeAgent의 유연한 코드 실행 방식에 구조화된 JSON 출력을 결합하면, 충분히 강한 모델에서 파싱 안정성과 추론 명시성이 높아져 여러 벤치마크 성능이 개선된다고 설명한다.

huggingface.co
#anthropic#agent-routing#llm#semiconductors
Shipping code faster with o3, o4-mini, and GPT-4.1
Article2025년 5월 22일

Shipping code faster with o3, o4-mini, and GPT-4.1

CodeRabbit은 코드 생성량이 아니라 리뷰 처리량이 실제 배포 속도를 제한한다는 문제의식에서 출발해, 저장소 맥락을 보강한 다단계 AI 리뷰로 정확하고 신속한 코드 배포를 지원한다.

openai.com
#agent-deployment#agent-routing#llm#semiconductors
New tools and features in the Responses API
Article2025년 5월 21일

New tools and features in the Responses API

OpenAI는 Responses API에 원격 MCP, 이미지 생성, 코드 인터프리터 등 새 도구와 장기 작업·추론 관리 기능을 추가해 에이전트 애플리케이션의 활용성, 신뢰성, 가시성, 개인정보 보호를 강화했다.

openai.com
#privacy-design#ai-safety#llm#semiconductors
The Transformers Library: standardizing model definitions
Article2025년 5월 15일

The Transformers Library: standardizing model definitions

트랜스포머스는 모델 정의를 표준화해 하나의 아키텍처 구현이 학습·추론·배포·로컬 실행 도구 전반으로 빠르게 이어지는 생태계의 중심축이 되고자 한다.

huggingface.co
#ai-architecture#llm#semiconductors#applications
LeRobot Community Datasets: The “ImageNet” of Robotics — When and How?
Article2025년 5월 11일

LeRobot Community Datasets: The “ImageNet” of Robotics — When and How?

로봇의 범용화는 모델 구조만의 문제가 아니라 다양한 환경·과제·기체에서 수집된 고품질 데이터를 함께 학습하는 문제이며, LeRobot 공동체는 개방형 로봇 데이터 생태계를 통해 로봇공학의 ‘이미지넷’을 만들어 가고 있다.

huggingface.co
#ai-architecture#multimodal#lerobot#llm
Argument Validation without Repetition
Article2025년 5월 5일

Argument Validation without Repetition

Convex와 TypeScript 프로젝트에서 스키마 검증기와 타입 유틸리티를 재사용해 인자 검증 중복을 줄이고, 함수·클라이언트·부분 업데이트까지 일관된 타입 안전성을 유지하는 방법을 설명한다.

stack.convex.dev
#agent-routing#llm#semiconductors#applications
Expanding on what we missed with sycophancy
Article2025년 5월 2일

Expanding on what we missed with sycophancy

오픈에이아이는 4월 25일 GPT 4o 업데이트에서 사용자 선호 지표와 여러 개선 요소가 결합해 과도한 동조 성향을 키웠으며, 기존 평가 체계가 이를 포착하지 못한 책임을 인정하고 롤백과 평가 절차 보강에 나섰다.

openai.com
#agent-deployment#agent-memory#context-compression#retrieval-index
Falcon-Edge: A series of powerful, universal, fine-tunable 1.58bit language models.
Article2025년 5월 1일

Falcon-Edge: A series of powerful, universal, fine-tunable 1.58bit language models.

Falcon Edge는 하나의 사전학습 과정에서 범용 bfloat16 모델, 삼진 가중치 기반 BitNet 모델, 미세조정용 사전 양자화 모델을 함께 제공하는 1.58비트 언어 모델 시리즈다.

huggingface.co
#ai-architecture#multimodal#llm#vision-language-models
Welcoming Llama Guard 4 on Hugging Face Hub
Article2025년 4월 29일

Welcoming Llama Guard 4 on Hugging Face Hub

메타가 공개한 라마 가드 4는 텍스트와 이미지를 함께 검사해 유해한 입력과 출력을 분류하는 120억 매개변수의 다국어 안전 모델이며, 프롬프트 주입과 탈옥을 탐지하는 라마 프롬프트 가드 2도 함께 제공된다.

huggingface.co
#privacy-design#ai-architecture#multimodal#agent-memory
Introducing our latest image generation model in the API
Article2025년 4월 23일

Introducing our latest image generation model in the API

OpenAI는 ChatGPT에서 인기를 얻은 이미지 생성 모델을 gpt image 1 API로 제공해, 개발자와 기업이 고품질 이미지 생성·편집 기능을 제품에 직접 통합할 수 있게 했다고 발표했다.

openai.com
#multimodal#ai-safety#llm#semiconductors
How Botpress Populates AI Chatbot Knowledge Bases at Scale with Firecrawl
Article2025년 4월 21일

How Botpress Populates AI Chatbot Knowledge Bases at Scale with Firecrawl

Botpress는 Firecrawl을 도입해 웹사이트 콘텐츠를 챗봇 지식 베이스로 가져오는 과정을 자동화하고, 자체 HTML 마크다운 처리 부담을 크게 줄였다.

Eric Ciarla
#agent-deployment#agent-routing#llm#semiconductors
LLM Inference on Edge: A Fun and Easy Guide to run LLMs via React Native on your Phone!
Article2025년 4월 21일

LLM Inference on Edge: A Fun and Easy Guide to run LLMs via React Native on your Phone!

이 글은 React Native 앱에서 Hugging Face의 GGUF 모델을 내려받고 llama.rn으로 로컬 실행해, Android와 iOS에서 오프라인 LLM 채팅 앱을 만드는 과정을 안내한다.

huggingface.co
#privacy-design#service-design#llm#semiconductors
How to think about agent frameworks
Article2025년 4월 20일

How to think about agent frameworks

제공된 source body는 ‘에이전트 프레임워크를 어떻게 생각할 것인가’라는 본문이 아니라, Webflow 기반 페이지의 전역 CSS·반응형 레이아웃·표·블로그 카드·리치 텍스트 스타일을 정의한 스타일 코드입니다.

LangChain Accounts
#anthropic#agent-routing#workflow-automation#travel-hospitality
Introducing HELMET: Holistically Evaluating Long-context Language Models
Article2025년 4월 20일

Introducing HELMET: Holistically Evaluating Long-context Language Models

HELMET은 장문맥 언어 모델을 실제 응용 과제에서 다양성·통제 가능성·신뢰성을 기준으로 종합 평가하며, 단순 합성 과제만으로는 드러나지 않는 모델별 강점과 장문 입력에서의 성능 저하를 밝히는 벤치마크다.

huggingface.co
#anthropic#long-context#llm#semiconductors
Announcing FIRE-1, Our Web Action Agent: Launch Week III - Day 2
Article2025년 4월 15일

Announcing FIRE-1, Our Web Action Agent: Launch Week III - Day 2

Firecrawl은 Launch Week III 2일차에 복잡한 웹사이트를 탐색하고 버튼·검색폼 등과 상호작용해 숨은 데이터를 추출하는 웹 액션 에이전트 FIRE 1을 발표했다.

Eric Ciarla
#agent-routing#context-compression#prompt-library#llm
Hugging Face to sell open-source robots thanks to Pollen Robotics acquisition 🤖
Article2025년 4월 14일

Hugging Face to sell open-source robots thanks to Pollen Robotics acquisition 🤖

허깅페이스는 오픈소스 휴머노이드 로봇 기업 폴렌 로보틱스를 인수하고, 르로봇 생태계와 리치 2를 결합해 개방형 로봇 소프트웨어·하드웨어를 직접 제공하기 시작했다.

huggingface.co
#multimodal#agent-routing#llm#semiconductors
Introducing Change Tracking: Launch Week III - Day 1
Article2025년 4월 14일

Introducing Change Tracking: Launch Week III - Day 1

Firecrawl은 웹사이트 스크랩·크롤 결과를 이전 버전과 비교해 신규, 동일, 변경, 제거 상태를 알려주는 Change Tracking 기능을 베타로 공개했다.

Eric Ciarla
#agent-routing#llm#semiconductors#applications
Get your VLM running in 3 simple steps on Intel CPUs
Article2025년 4월 8일

Get your VLM running in 3 simple steps on Intel CPUs

소형 비전 언어 모델 SmolVLM2를 Optimum Intel과 OpenVINO로 변환·양자화·추론하면 별도 GPU 없이도 Intel CPU에서 지연시간과 처리량을 크게 개선할 수 있다.

huggingface.co
#privacy-design#service-design#multimodal#agent-memory
🚀 Accelerating LLM Inference with TGI on Intel Gaudi
Article2025년 3월 29일

🚀 Accelerating LLM Inference with TGI on Intel Gaudi

Hugging Face는 Intel Gaudi 하드웨어 지원을 Text Generation Inference(TGI) 본체에 네이티브로 통합해, 별도 포크 없이 Gaudi 기반 LLM 추론 배포를 사용할 수 있게 했습니다.

huggingface.co
#ai-architecture#multimodal#llm#semiconductors
How to deploy and fine-tune DeepSeek models on AWS
Article2025년 3월 27일

How to deploy and fine-tune DeepSeek models on AWS

이 글은 Hugging Face 도구를 이용해 DeepSeek R1 계열 모델을 여러 환경에 배포하는 방법과 모델별 하드웨어 구성, 현재 지원되는 기능과 아직 준비 중인 미지원 영역을 정리한 실무 가이드다.

huggingface.co
#agent-deployment#ai-infrastructure#capex-cycle#llm
Mixture of Experts (MoEs) in Transformers
Article2025년 3월 27일

Mixture of Experts (MoEs) in Transformers

전문가 혼합 모델은 토큰마다 일부 전문가만 활성화해 전체 모델 용량과 실제 계산량을 분리하며, Transformers는 이를 효율적으로 지원하기 위해 가중치 로딩·실행 백엔드·전문가 병렬화 구조를 재설계했다.

huggingface.co
#kimi-k2-5#ai-architecture#transformer#agent-memory
Merging Streams of Convex data
Article2025년 3월 21일

Merging Streams of Convex data

이 글은 Convex에서 convex helpers의 스트림 기능을 사용해 여러 인덱스 범위의 문서를 병합하고, 조인형 확장과 변환, 필터링, 페이지네이션을 더 효율적으로 구현하는 방법을 설명한다.

stack.convex.dev
#agent-routing#llm#semiconductors#applications
이전1…18192021222320 / 23다음