Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#long-context
Tag47건Article 47

#long-context

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#mamba공동문서 2 · 연관도 21%#llm공동문서 43 · 연관도 17%#ai-architecture공동문서 22 · 연관도 15%#agent-orchestration-reliability공동문서 1 · 연관도 15%#agentic-memory공동문서 1 · 연관도 15%#agentic-rl공동문서 1 · 연관도 15%#agentic-rl-debugging공동문서 1 · 연관도 15%#agentic-science-tools공동문서 1 · 연관도 15%#ai-xr-prototyping공동문서 1 · 연관도 15%#android-xr공동문서 1 · 연관도 15%
Titans + MIRAS: Helping AI have long-term memory
Article2025년 12월 4일

Titans + MIRAS: Helping AI have long-term memory

Titans와 MIRAS는 실행 중 장기 기억을 선택적으로 갱신하는 방식으로, 매우 긴 문맥을 더 빠르고 정확하게 다루려는 새로운 시퀀스 모델링 접근이다.

research.google
#ai-architecture#agent-memory#context-compression#retrieval-index
Introducing GPT-5.1 for developers
Article2025년 11월 13일

Introducing GPT-5.1 for developers

GPT 5.1은 작업 난도에 따라 추론량을 조절하고, 추론 없는 저지연 모드와 24시간 프롬프트 캐싱, 향상된 코딩 성능, 코드 패치 및 셸 도구를 제공하는 개발자용 모델이다.

openai.com
#token-efficiency#agent-routing#prompt-library#workflow-automation
Introducing Nested Learning: A new ML paradigm for continual learning
Article2025년 11월 7일

Introducing Nested Learning: A new ML paradigm for continual learning

구글 리서치는 모델 구조와 최적화 규칙을 하나의 중첩된 학습 시스템으로 보는 Nested Learning을 제안하며, 지속 학습의 핵심 난제인 파국적 망각을 줄이는 새 설계 관점을 제시했다.

research.google
#privacy-design#ai-architecture#agent-memory#agent-routing
Introducing the Open Leaderboard for Japanese LLMs!
Article2025년 10월 20일

Introducing the Open Leaderboard for Japanese LLMs!

LLM jp와 허깅페이스가 일본어의 복합적인 언어 특성을 반영한 16개 과제와 20개 이상의 데이터셋으로 개방형 일본어 대규모 언어 모델을 비교·분석하는 공개 리더보드를 구축했다.

huggingface.co
#service-design#ai-architecture#context-compression#prompt-library
A picture's worth a thousand (private) words: Hierarchical generation of coherent synthetic photo albums
Article2025년 10월 20일

A picture's worth a thousand (private) words: Hierarchical generation of coherent synthetic photo albums

Google Research는 사진 앨범을 먼저 텍스트 요약과 사진별 캡션으로 바꾼 뒤 차등 개인정보보호를 적용해 계층적으로 생성함으로써, 개인정보를 보호하면서도 주제 일관성을 유지하는 합성 사진 앨범 생성 방법을 제안했다.

research.google
#privacy-design#agent-routing#yfcc100m-mauve#llm
mmBERT: ModernBERT goes Multilingual
Article2025년 10월 7일

mmBERT: ModernBERT goes Multilingual

mmBERT는 1,833개 언어와 3조 개 이상의 토큰으로 학습해 XLM R을 넘어서는 다국어 성능, 최대 8,192토큰의 문맥 처리, 기존 다국어 인코더 대비 2~4배의 효율을 함께 달성한 ModernBERT 기반 인코더 모델이다.

huggingface.co
#ai-architecture#llm#applications#gpu
Introducing RTEB: A New Standard for Retrieval Evaluation
Article2025년 10월 1일

Introducing RTEB: A New Standard for Retrieval Evaluation

RTEB는 공개·비공개 데이터셋을 함께 활용해 임베딩 모델의 미지 데이터 검색 성능과 실제 응용 적합성을 공정하게 평가하려는 새로운 검색 벤치마크다.

huggingface.co
#multimodal#agent-memory#change-management#organizational-redesign
Time series foundation models can be few-shot learners
Article2025년 9월 23일

Time series foundation models can be few-shot learners

구글 리서치는 TimesFM에 지속 사전학습과 구분 토큰을 더해, 사용자가 별도 지도 미세조정을 하지 않아도 관련 시계열 예시 몇 개를 추론 시점에 활용하는 TimesFM ICF를 제안했다.

research.google
#ai-architecture#llm#semiconductors#applications
TextQuests: How Good are LLMs at Text-Based Video Games?
Article2025년 8월 13일

TextQuests: How Good are LLMs at Text-Based Video Games?

TextQuests는 25개의 고전 텍스트 어드벤처 게임을 통해 LLM 에이전트의 장기 문맥 추론, 탐색 학습, 공간 이해, 행동 효율성과 유해 행동 경향을 평가하는 벤치마크다.

huggingface.co
#token-efficiency#llm#semiconductors#applications
Introducing GPT‑5 for developers
Article2025년 8월 7일

Introducing GPT‑5 for developers

GPT‑5는 코딩 정확도와 도구 호출, 지시 이행, 장문 맥락 검색을 함께 개선하고 개발자가 응답 방식과 추론 수준을 조절할 수 있게 설계된 OpenAI의 API용 추론 모델이다.

openai.com
#openai#long-context#privacy-design#service-design
Benchmarking Language Model Performance on 5th Gen Xeon at GCP
Article2025년 7월 29일

Benchmarking Language Model Performance on 5th Gen Xeon at GCP

Google Cloud의 5세대 Xeon 기반 C4 인스턴스는 3세대 Xeon 기반 N2보다 텍스트 임베딩과 텍스트 생성에서 큰 처리량 및 비용 효율 우위를 보이며, 경량 에이전틱 AI를 CPU만으로 배포할 가능성을 보여준다.

huggingface.co
#llm#semiconductors#applications#long-context
Jupyter Agents: training LLMs to reason with notebooks
Article2025년 7월 26일

Jupyter Agents: training LLMs to reason with notebooks

Jupyter Agent 프로젝트는 노트북 안에서 코드 실행과 추론을 결합하는 데이터 과학 에이전트를 만들고, Kaggle 노트북 기반 데이터 파이프라인으로 소형 Qwen3 4B 모델의 DABStep 성능을 끌어올리려는 시도다.

huggingface.co
#hugging-face#jupyter-agent#qwen3-32b#qwen-3-coder
Ulysses Sequence Parallelism: Training with Million-Token Contexts
Article2025년 7월 26일

Ulysses Sequence Parallelism: Training with Million-Token Contexts

Ulysses Sequence Parallelism은 긴 시퀀스 학습에서 시퀀스와 어텐션 헤드를 함께 나누고 all to all 통신으로 재배치해, 단일 GPU 메모리 한계를 넘어 수십만~백만 토큰 문맥 학습을 가능하게 하는 방식이다.

huggingface.co
#ai-architecture#agent-memory#agent-routing#retrieval-index
TimeScope: How Long Can Your Video Large Multimodal Model Go?
Article2025년 7월 23일

TimeScope: How Long Can Your Video Large Multimodal Model Go?

TimeScope는 1분부터 8시간까지의 영상에 짧은 동영상 클립을 삽입해 검색·정보 종합·세밀한 시간 지각 능력을 측정하며, 최신 비전 언어 모델의 장시간 영상 이해가 아직 제한적임을 보여주는 오픈소스 벤치마크다.

huggingface.co
#multimodal#llm#semiconductors#vision-language-models
Ettin Suite: SoTA Paired Encoders and Decoders
Article2025년 7월 16일

Ettin Suite: SoTA Paired Encoders and Decoders

에틴은 동일한 공개 데이터 2조 토큰, 모델 구조, 학습 절차를 적용한 1,700만~10억 매개변수 규모의 인코더·디코더 쌍으로, 두 아키텍처를 공정하게 비교하면서 양쪽 모두에서 공개 데이터 기반 최고 수준의 성능을 달성한 모델군이다.

huggingface.co
#ai-architecture#llm#semiconductors#applications
Introducing Falcon-H1-Arabic: Pushing the Boundaries of Arabic Language AI with Hybrid Architecture
Article2025년 7월 14일

Introducing Falcon-H1-Arabic: Pushing the Boundaries of Arabic Language AI with Hybrid Architecture

팔콘 에이치원 아라빅은 맘바와 트랜스포머를 결합한 하이브리드 구조, 최대 25만 6천 토큰의 문맥 창, 아랍어 특화 데이터와 후속 학습을 통해 규모별 최고 수준의 아랍어 처리 성능을 목표로 한 3종 모델군이다.

huggingface.co
#ai-architecture#agent-routing#workflow-automation#llm
SmolLM3: smol, multilingual, long-context reasoner
Article2025년 7월 8일

SmolLM3: smol, multilingual, long-context reasoner

SmolLM3는 공개 데이터와 학습 도구로 구축한 30억 매개변수 모델로, 11조 개가 넘는 토큰의 단계별 사전학습과 장문·추론 중간학습을 결합해 다국어, 최대 128K 문맥, 추론·비추론 이중 모드를 지원한다.

huggingface.co
#long-context#privacy-design#ai-architecture#llm
Welcome Llama 4 Maverick & Scout on Hugging Face
Article2025년 5월 22일

Welcome Llama 4 Maverick & Scout on Hugging Face

Meta의 네이티브 멀티모달 MoE 모델 Llama 4 Maverick과 Scout가 출시 당일부터 Hugging Face Hub, Transformers, TGI, 양자화 및 Xet 저장소를 통해 제공되며, 최대 1천만 토큰 문맥과 강력한 추론·이미지·코딩 성능을 지원한다.

huggingface.co
#ai-architecture#multimodal#agent-deployment#agent-routing
Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models
Article2025년 4월 30일

Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models

이 글은 Qwen3 8B를 Intel® Core™ Ultra에서 더 빠르게 실행하기 위해 OpenVINO.GenAI의 추측 디코딩과 깊이 가지치기된 Qwen3 0.6B 드래프트 모델을 결합해 약 1.4배 속도 향상을 얻은 과정을 설명한다.

huggingface.co
#service-design#agent-routing#capex-cycle#context-compression
Introducing HELMET: Holistically Evaluating Long-context Language Models
Article2025년 4월 20일

Introducing HELMET: Holistically Evaluating Long-context Language Models

HELMET은 장문맥 언어 모델을 실제 응용 과제에서 다양성·통제 가능성·신뢰성을 기준으로 종합 평가하며, 단순 합성 과제만으로는 드러나지 않는 모델별 강점과 장문 입력에서의 성능 저하를 밝히는 벤치마크다.

huggingface.co
#anthropic#long-context#llm#semiconductors
Open-R1: a fully open reproduction of DeepSeek-R1
Article2025년 3월 27일

Open-R1: a fully open reproduction of DeepSeek-R1

Open R1은 DeepSeek R1에서 공개되지 않은 데이터셋과 학습 코드를 재구성해, 추론 모델의 증류·순수 강화학습·다단계 학습 과정을 누구나 검증하고 재현할 수 있도록 만들려는 오픈 프로젝트다.

huggingface.co
#anthropic#agent-routing#workflow-automation#llm
AI and the Future of Cybersecurity: Why Openness Matters
Article2025년 2월 4일

AI and the Future of Cybersecurity: Why Openness Matters

이 글은 Mythos 사례를 통해 AI 사이버보안의 핵심이 단일 모델이 아니라 시스템·생태계에 있으며, 방어자가 공격자와 맞서기 위해서는 개방형 도구와 감사 가능한 구조가 중요하다고 설명한다.

huggingface.co
#llm#semiconductors#ai-coding#long-context
Open R1: How to use OlympicCoder locally for coding
Article2025년 1월 12일

Open R1: How to use OlympicCoder locally for coding

올림픽코더 7B의 4비트 양자화 모델을 엘엠 스튜디오에서 구동하고 컨티뉴 확장 기능으로 비주얼 스튜디오 코드에 연결해 로컬 코딩 도우미로 사용하는 방법을 설명한다.

huggingface.co
#llm#semiconductors#applications#long-context
이전122 / 2다음