Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#long-context
Tag47건Article 47

#long-context

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#mamba공동문서 2 · 연관도 21%#llm공동문서 43 · 연관도 17%#ai-architecture공동문서 22 · 연관도 15%#agent-orchestration-reliability공동문서 1 · 연관도 15%#agentic-memory공동문서 1 · 연관도 15%#agentic-rl공동문서 1 · 연관도 15%#agentic-rl-debugging공동문서 1 · 연관도 15%#agentic-science-tools공동문서 1 · 연관도 15%#ai-xr-prototyping공동문서 1 · 연관도 15%#android-xr공동문서 1 · 연관도 15%
Run MiniMax models on Amazon Bedrock
Article2026년 7월 6일

Run MiniMax models on Amazon Bedrock

이 글은 Amazon Bedrock에서 MiniMax M2 계열 모델을 사용하는 방법과 모델 선택 기준, 접근 엔드포인트, 초기 설정 및 도구 호출 예제를 설명한다.

aws.amazon.com
#service-design#ai-architecture#multimodal#agent-deployment
How to Use RLMs in Deep Agents
Article2026년 7월 1일

How to Use RLMs in Deep Agents

Deep Agents는 RLM에서 영감을 받은 동적 서브에이전트와 코드 인터프리터를 통해 긴 컨텍스트 작업을 턴별 추론이 아니라 프로그래밍 가능한 재귀적 워크플로로 처리하려 한다.

langchain.com
#ai-architecture#agent-routing#workflow-automation#llm
The Agent Development Lifecycle: Build, Test, Deploy & Monitor AI Agents
Article2026년 6월 25일

The Agent Development Lifecycle: Build, Test, Deploy & Monitor AI Agents

이 글은 AI 에이전트를 일회성 데모가 아니라 반복적으로 구축·검증·배포·관찰하며 개선하는 ‘Agent Development Lifecycle’을 Build, Test, Deploy, Monitor 네 단계로 설명한다.

langchain.com
#agent-deployment#agent-routing#context-compression#prompt-library
A startup claims it broke through a bottleneck that’s holding back LLMs
Article2026년 6월 19일

A startup claims it broke through a bottleneck that’s holding back LLMs

마이애미 기반 스타트업 Subquadratic은 LLM의 핵심 병목으로 지목돼 온 dense attention의 계산 비용 문제를 sparse attention 방식의 SubQ로 완화했다고 주장하며, Appen의 일부 독립 평가 결과가 그 주장에 주목할 만한 근거를 제공했다.

technologyreview.com
#anthropic#ai-architecture#llm#semiconductors
GLM-5.2: Built for Long-Horizon Tasks
Article2026년 6월 17일

GLM-5.2: Built for Long-Horizon Tasks

GLM 5.2는 안정적인 100만 토큰 문맥, 강화된 코딩 성능, 효율적인 장문맥 구조와 서빙·강화학습 체계를 결합해 장시간 수행되는 복잡한 엔지니어링 작업을 겨냥한 공개형 주력 모델이다.

huggingface.co
#service-design#ai-architecture#agent-routing#llm
Introducing North Mini Code: Cohere’s First Model For Developers
Article2026년 6월 15일

Introducing North Mini Code: Cohere’s First Model For Developers

코히어가 에이전트형 소프트웨어 엔지니어링에 특화된 300억 매개변수 희소 전문가 혼합 모델 ‘노스 미니 코드’를 아파치 2.0 라이선스로 공개했다.

huggingface.co
#ai-architecture#agent-routing#llm#semiconductors
DeepSeek-V4: a million-token context that agents can actually use
Article2026년 6월 8일

DeepSeek-V4: a million-token context that agents can actually use

DeepSeek V4는 최고 벤치마크 점수보다 100만 토큰 문맥을 실제 에이전트 작업에서 감당하게 만드는 긴 문맥 효율, 도구 호출 지속성, 샌드박스 기반 학습 인프라에 초점을 둔 모델이다.

huggingface.co
#ai-architecture#agent-memory#agent-routing#context-compression
A New Era of Discovery: Google Research at I/O 2026
Article2026년 5월 28일

A New Era of Discovery: Google Research at I/O 2026

Google Research는 I/O 2026에서 Gemini for Science, 의료 AI, 엣지 AI 하드웨어, 재난 예측 모델을 통해 AI가 과학·보건·기후 대응의 실제 발견과 의사결정을 가속하는 방향을 제시했다.

Google
#alphaevolve#co-scientist#google-research#gemini-for-science
Introducing the Ettin Reranker Family
Article2026년 5월 19일

Introducing the Ettin Reranker Family

Ettin Reranker는 1,760만~10억 매개변수의 여섯 가지 교차 인코더로, 검색 속도와 재정렬 품질 사이에서 선택 가능한 장문 지원 모델군이다.

huggingface.co
#service-design#ai-architecture#llm#ai-coding
Databricks brings GPT-5.5 to enterprise agent workflows
Article2026년 5월 15일

Databricks brings GPT-5.5 to enterprise agent workflows

Databricks는 복잡한 기업 문서 작업 벤치마크인 OfficeQA Pro에서 GPT 5.5가 새 최고 성능을 기록하자, 이를 고객용 엔터프라이즈 에이전트 워크플로에 제공하기 시작했다.

openai.com
#databricks#openai#gpt-5-4#gpt-5-5
Top Companies Are Secretly Working on This (It Will Replace LLMs)
Article2026년 5월 5일

Top Companies Are Secretly Working on This (It Will Replace LLMs)

SSM은 긴 컨텍스트에서 Transformer의 비용·메모리 병목을 줄이기 위한 대안적 시퀀스 처리 구조로, 특히 장기 작업을 수행하는 에이전트 시스템에서 주목받고 있다는 것이 원문의 핵심 주장입니다.

Siddharth
#hyena#mamba#mamba-2#state-space-model
The PR you would have opened yourself
Article2026년 4월 16일

The PR you would have opened yourself

transformers 모델을 mlx lm으로 빠르고 정확하게 이식하도록 돕는 스킬과 비에이전트 테스트 하네스를 구축하되, 코드 소유권과 최종 판단은 기여자와 리뷰어에게 남겨 두는 접근을 설명한다.

huggingface.co
#privacy-design#ai-architecture#multimodal#llm
Training mRNA Language Models Across 25 Species for $165
Article2026년 3월 31일

Training mRNA Language Models Across 25 Species for $165

OpenMed는 단백질 구조 예측, 서열 설계, 코돈 최적화를 잇는 단백질 AI 파이프라인을 구축하고, 코돈 수준 언어모델 비교 끝에 CodonRoBERTa large v2가 생물학적 지표에서 가장 유용하다는 결론을 제시했다.

huggingface.co
#ai-architecture#agent-routing#workflow-automation#llm
Vibe Coding XR: Accelerating AI + XR prototyping with XR Blocks and Gemini
Article2026년 3월 25일

Vibe Coding XR: Accelerating AI + XR prototyping with XR Blocks and Gemini

Vibe Coding XR는 Gemini와 오픈소스 XR Blocks를 결합해 자연어 프롬프트를 60초 안팎의 상호작용형 WebXR/Android XR 프로토타입으로 바꾸는 빠른 XR 제작 워크플로입니다.

research.google
#gemini#gemini-canvas#google-xr#xr-blocks
TurboQuant: Redefining AI efficiency with extreme compression
Article2026년 3월 24일

TurboQuant: Redefining AI efficiency with extreme compression

TurboQuant는 QJL과 PolarQuant를 결합해 벡터 양자화의 메모리 오버헤드를 줄이고, KV 캐시 압축과 고차원 벡터 검색에서 정확도 손실 없이 큰 압축·속도 향상을 보인 Google Research의 알고리즘입니다.

research.google
#privacy-design#agent-memory#context-compression#retrieval-index
Ecom-RLVE: Adaptive Verifiable Environments for E-Commerce Conversational Agents
Article2026년 3월 8일

Ecom-RLVE: Adaptive Verifiable Environments for E-Commerce Conversational Agents

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

huggingface.co
#long-context#service-design#agent-deployment#agent-routing
Gemma 3n fully available in the open-source ecosystem!
Article2026년 3월 6일

Gemma 3n fully available in the open-source ecosystem!

젬마 3n은 적은 GPU 메모리로 구동되는 네이티브 멀티모달 모델로, 트랜스포머스·MLX·라마닷씨피피·트랜스포머스닷제이에스 등 주요 오픈소스 생태계에서 추론과 미세조정을 지원한다.

huggingface.co
#ai-architecture#multimodal#agent-memory#agent-routing
Introducing GPT-5.4
Article2026년 3월 5일

Introducing GPT-5.4

OpenAI는 GPT 5.4를 ChatGPT, API, Codex에 공개하며 전문 업무, 코딩, 컴퓨터 사용 에이전트, 장기 작업 효율성을 강화했다고 밝혔다.

openai.com
#service-design#token-efficiency#agent-routing#llm
GGML and llama.cpp join HF to ensure the long-term progress of Local AI
Article2026년 2월 20일

GGML and llama.cpp join HF to ensure the long-term progress of Local AI

GGML·llama.cpp 팀이 허깅페이스에 합류해 기술적 자율성과 오픈소스 운영을 유지하면서 로컬 AI 생태계의 지속 가능성, 모델 지원 속도, 사용자 접근성을 함께 강화한다.

huggingface.co
#agent-deployment#llm#semiconductors#applications
Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents
Article2026년 1월 30일

Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents

NVIDIA Nemotron 3 Nano Omni는 문서, 이미지, 비디오, 오디오를 긴 컨텍스트 안에서 함께 이해하도록 설계된 오픈 웨이트 옴니모달 모델로, 문서 지능·영상/음성 이해·GUI 에이전트 작업에서 높은 정확도와 효율을 내세운다.

huggingface.co
#mamba#nvidia#hugging-face#c-radiov4-h
Unlocking Agentic RL Training for GPT-OSS: A Practical Retrospective
Article2026년 1월 27일

Unlocking Agentic RL Training for GPT-OSS: A Practical Retrospective

이 글은 GPT OSS를 에이전트형 강화학습의 백본 모델로 활용하기 위해 verl 기반 PPO 학습에서 발견한 온폴리시 불일치, 훈련·추론 불일치, attention sink 미지원 문제를 단계적으로 진단하고 수정한 실험 회고다.

huggingface.co
#gsm8k#retool#gpt-oss-120b#gpt-oss-20b
Differential Transformer V2
Article2026년 1월 20일

Differential Transformer V2

DIFF V2는 같은 GQA 그룹의 두 쿼리 헤드 출력을 토큰·헤드별 계수로 차감해, 키·값 헤드와 표준 어텐션 커널을 그대로 유지하면서 디코딩 효율, 학습 안정성, 표현 자유도, 출력 투영의 매개변수 효율을 함께 개선한 구조다.

huggingface.co
#ai-architecture#multimodal#llm#semiconductors
Introducing GPT-5.2-Codex
Article2025년 12월 18일

Introducing GPT-5.2-Codex

GPT‑5.2‑Codex는 장기적이고 복잡한 소프트웨어 엔지니어링과 방어적 보안 연구 역량을 강화하면서, 이중용도 위험에 대응하기 위한 보호 장치와 단계적 접근 정책을 함께 도입한 에이전트형 코딩 모델이다.

openai.com
#service-design#token-efficiency#agent-routing#llm
DeepMath: A lightweight math reasoning Agent with smolagents
Article2025년 12월 8일

DeepMath: A lightweight math reasoning Agent with smolagents

DeepMath는 Qwen3 4B Thinking 기반의 경량 수학 추론 에이전트로, 긴 사고 과정을 짧은 파이썬 실행 단계로 대체하고 GRPO 학습을 통해 더 간결하고 정확한 수학 풀이를 목표로 한다.

huggingface.co
#llm#semiconductors#applications#long-context
이전121 / 2다음