Uploaded June 2022 | Updated September 2026, 2 weeks ago
#openai #vpt #minecraft
Minecraft is one of the harder challenges any RL agent could face. Episodes are long, and the world is procedurally generated, complex, and huge. Further, the action space is a keyboard and a mouse, which has to be operated only given the game's video input. OpenAI tackles this challenge using Video PreTraining, leveraging a small set of contractor data in order to pseudo-label a giant corpus of scraped footage of gameplay. The pre-trained model is highly capable in basic game mechanics and can be fine-tuned much better than a blank slate model. This is the first Minecraft agent that achieves the elusive goal of crafting a diamond pickaxe all by itself.
OUTLINE:
0:00 - Intro
3:50 - How to spend money most effectively?
8:20 - Getting a large dataset with labels
14:40 - Model architecture
19:20 - Experimental results and fine-tuning
25:40 - Reinforcement Learning to the Diamond Pickaxe
30:00 - Final comments and hardware
Blog: openai.com/blog/vpt
Paper: arxiv.org/abs/2206.11795
Code & Model weights: github.com/openai/Video-Pre-Training
Abstract:
Pretraining on noisy, internet-scale datasets has been heavily studied as a technique for training models with broad, general capabilities for text, images, and other modalities. However, for many sequential decision domains such as robotics, video games, and computer use, publicly available data does not contain the labels required to train behavioral priors in the same way. We extend the internet-scale pretraining paradigm to sequential decision domains through semi-supervised imitation learning wherein agents learn to act by watching online unlabeled videos. Specifically, we show that with a small amount of labeled data we can train an inverse dynamics model accurate enough to label a huge unlabeled source of online data -- here, online videos of people playing Minecraft -- from which we can then train a general behavioral prior. Despite using the native human interface (mouse and keyboard at 20Hz), we show that this behavioral prior has nontrivial zero-shot capabilities and that it can be fine-tuned, with both imitation learning and reinforcement learning, to hard-exploration tasks that are impossible to learn from scratch via reinforcement learning. For many tasks our models exhibit human-level performance, and we are the first to report computer agents that can craft diamond tools, which can take proficient humans upwards of 20 minutes (24,000 environment actions) of gameplay to accomplish.
Authors: Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, Jeff Clune
Links:
Homepage: ykilcher.com
Merch: ykilcher.com/merch
YouTube: youtube.com/c/yannickilcher
Twitter: twitter.com/ykilcher
Discord: ykilcher.com/discord
LinkedIn: linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: subscribestar.com/yannickilcher
Patreon: patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
#openai #vpt #minecraft
Minecraft is one of the harder challenges any RL agent could face. Episodes are long, and the world is procedurally generated, complex, and huge. Further, the action space is a keyboard and a mouse, which has to be operated only given the game's video input. OpenAI tackles this challenge using Video PreTraining, leveraging a small set of contractor data in order to pseudo-label a giant corpus of scraped footage of gameplay. The pre-trained model is highly capable in basic game mechanics and can be fine-tuned much better than a blank slate model. This is the first Minecraft agent that achieves the elusive goal of crafting a diamond pickaxe all by itself.
OUTLINE:
0:00 - Intro
3:50 - How to spend money most effectively?
8:20 - Getting a large dataset with labels
14:40 - Model architecture
19:20 - Experimental results and fine-tuning
25:40 - Reinforcement Learning to the Diamond Pickaxe
30:00 - Final comments and hardware
Blog: openai.com/blog/vpt
Paper: arxiv.org/abs/2206.11795
Code & Model weights: github.com/openai/Video-Pre-Training
Abstract:
Pretraining on noisy, internet-scale datasets has been heavily studied as a technique for training models with broad, general capabilities for text, images, and other modalities. However, for many sequential decision domains such as robotics, video games, and computer use, publicly available data does not contain the labels required to train behavioral priors in the same way. We extend the internet-scale pretraining paradigm to sequential decision domains through semi-supervised imitation learning wherein agents learn to act by watching online unlabeled videos. Specifically, we show that with a small amount of labeled data we can train an inverse dynamics model accurate enough to label a huge unlabeled source of online data -- here, online videos of people playing Minecraft -- from which we can then train a general behavioral prior. Despite using the native human interface (mouse and keyboard at 20Hz), we show that this behavioral prior has nontrivial zero-shot capabilities and that it can be fine-tuned, with both imitation learning and reinforcement learning, to hard-exploration tasks that are impossible to learn from scratch via reinforcement learning. For many tasks our models exhibit human-level performance, and we are the first to report computer agents that can craft diamond tools, which can take proficient humans upwards of 20 minutes (24,000 environment actions) of gameplay to accomplish.
Authors: Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, Jeff Clune
Links:
Homepage: ykilcher.com
Merch: ykilcher.com/merch
YouTube: youtube.com/c/yannickilcher
Twitter: twitter.com/ykilcher
Discord: ykilcher.com/discord
LinkedIn: linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: subscribestar.com/yannickilcher
Patreon: patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n


![[ML News] Metas OPT 175B language model | DALL-E Mega is training | TorToiSe TTS fakes my voice
#mlnews #dalle #gpt3
An inside look of whats happening in the ML world!
Sponsor: Weights & Biases
https://wandb.me/yannic
OUTLINE:
0:00 - Intro
0:20 - Sponsor: Weights & Biases
1:40 - Meta AI releases OPT-175B
4:55 - CoCa: New CLIP-Competitor
8:15 - DALL-E Mega is training
10:05 - TorToiSe TTS is amazing!
11:50 - Investigating Vision Transformers
12:50 - Hugging Face Deep RL class launched
13:40 - Helpful Things
17:00 - John Deeres driverless tractors
References:
Meta AI releases OPT-175B
https://ai.facebook.com/blog/democratizing-access-to-large-scale-language-models-with-opt-175b/
https://arxiv.org/abs/2205.01068
https://arxiv.org/pdf/2205.01068.pdf
https://github.com/facebookresearch/metaseq/tree/main/projects/OPT
https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/OPT175B_Logbook.pdf
https://github.com/facebookresearch/metaseq/tree/main/projects/OPT/chronicles
https://twitter.com/yoavgo/status/1522150063815987201
CoCa: New CLIP-Competitor
https://arxiv.org/abs/2205.01917
https://arxiv.org/pdf/2205.01917.pdf
DALL-E Mega is training
https://twitter.com/borisdayma
https://twitter.com/borisdayma/status/1521891895001112577
https://wandb.ai/dalle-mini/dalle-mini/reports/DALL-E-Mega VmlldzoxODMxMDI2
TorToiSe TTS is amazing!
https://github.com/neonbjb/tortoise-tts
https://nonint.com/static/tortoise_v2_examples.html
https://colab.research.google.com/drive/1wVVqUPqwiDBUVeWWOUNglpGhU3hg_cbR
https://github.com/neonbjb
Investigating Vision Transformers
https://github.com/sayakpaul/probing-vits/?utm_source=pocket_mylist
https://twitter.com/RisingSayak/status/1515918406171914240?utm_source=pocket_mylist
https://keras.io/examples/vision/probing_vits/
https://github.com/sayakpaul/probing-vits/tree/main/notebooks?utm_source=pocket_mylist
Hugging Face Deep RL class launched
https://github.com/huggingface/deep-rl-class
Helpful Things
https://merantix-momentum.com/technology/squirrel/?utm_source=pocket_mylist
https://github.com/merantix-momentum/squirrel-core?utm_source=pocket_mylist
https://pyscript.net/?utm_source=pocket_mylist
https://github.com/google-research/big_vision
https://deepsportradar.github.io/challenge.html
https://github.com/DeepSportRadar/camera-calibration-challenge
https://twitter.com/alekseykorshuk/status/1515989357961920514?utm_source=pocket_mylist
https://github.com/AlekseyKorshuk/huggingnft
John Deeres driverless tractors
https://thenextweb.com/news/john-deere-slowly-becoming-one-worlds-most-important-ai-companies
https://tractorhacking.github.io/
Links:
Merch: https://ykilcher.com/merch
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
BitChute: https://www.bitchute.com/channel/yannic-kilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n [ML News] Metas OPT 175B language model | DALL-E Mega is training | TorToiSe TTS fakes my voice](https://i.ytimg.com/vi/pwSnC8jlh50/mqdefault.jpg)
![[ML News] Devin AI Software Engineer | GPT-4.5-Turbo LEAKED | US Govt Report: Total Extinction
Your weekly dose of ML News
OUTLINE:
0:00 - Intro
0:15 - Devin: AI software engineer
5:50 - Mira Murati on Sora training data
6:50 - Inflection accused of copying Claude
9:00 - Tools & papers
16:30 - GPT-4.5-turbo mystery
17:30 - US government report: total extinction by AI
19:20 - Various other news
References:
https://www.cognition-labs.com/introducing-devin
https://twitter.com/cognition_labs/status/1767548763134964000?t=ZECIn-uqbguwHtY8X_Gvtw&s=09
https://news.google.com/stories/CAAqNggKIjBDQklTSGpvSmMzUnZjbmt0TXpZd1NoRUtEd2lWMUwyU0N4RnVWM3pSRWhWX01pZ0FQAQ?hl=en-US&gl=US&ceid=US%3Aen
https://www.bloomberg.com/news/articles/2024-03-12/cognition-ai-is-a-peter-thiel-backed-coding-assistant?embedded-checkout=true
https://www.bloomberg.com/authors/AQWHkoPod9g/ashlee-vance
https://www.bloomberg.com/news/articles/2024-03-12/cognition-ai-is-a-peter-thiel-backed-coding-assistant?srnd=undefined&embedded-checkout=true
https://www.bloomberg.com/news/newsletters/2024-03-12/cognition-ai-s-devin-assistant-can-build-websites-videos-from-a-prompt?srnd=undefined&embedded-checkout=true
https://archive.ph/5LZV9
https://github.com/opendevin/opendevin
https://twitter.com/MetaGPT_/status/1767965444579692832?t=dsYKmPfOBVGCFCwvPtZVWQ&s=09
https://docs.deepwisdom.ai/main/en/DataInterpreter/detail.html?id=AppleStockPriceAnalysisAndPrediction
https://docs.deepwisdom.ai/main/en/guide/use_cases/agent/interpreter/intro.html
https://github.com/geekan/MetaGPT/tree/main/examples/di
https://inflection.ai/inflection-2-5
https://twitter.com/seshubon/status/1765870717844050221
https://twitter.com/inflectionAI/status/1766173427441049684
https://www.mlxserver.com/
https://huggingface.co/spaces/mlabonne/AutoMerger
https://github.com/microsoft/aici
https://github.com/google-research/google-research/tree/master/fax
https://github.com/stanfordnlp/pyvene
https://arxiv.org/pdf/2403.06634.pdf
https://twitter.com/mattshumer_/status/1767606938538295757?t=1dYect5ylg9xrWSS4sL38Q&s=09
https://time.com/6898967/ai-extinction-national-security-risks-report/
https://venturebeat.com/ai/hugging-face-is-launching-an-open-source-robotics-project-led-by-former-tesla-scientist/
https://twitter.com/gcabanac/status/1767574447337124290?t=MnzwEbf_Zx0yQthe0RQ8hw&s=09
https://twitter.com/AnthropicAI/status/1768018310615151002?t=3ieMvNZxaoTXGGZttBsBvQ&s=09
https://huggingface.co/CohereForAI/c4ai-command-r-v01
https://twitter.com/Yampeleg/status/1765707714473197729?t=p3zOXUqKdqS-RzYjTNo65g&s=09
https://huggingface.co/yam-peleg/Hebrew-Gemma-11B
https://enriccorona.github.io/vlogger/
https://huggingface.co/NousResearch/Genstruct-7B
https://deepmind.google/discover/blog/sima-generalist-ai-agent-for-3d-virtual-environments/
https://arxiv.org/abs/2403.04652
https://twitter.com/corry_wang/status/1766949316394897851?t=i0ndsef_I_b3BDkVmyHgYw&s=09
https://twitter.com/sama/status/1766291001134715207?t=Wgyye9odOfF1Aoo0hZGihg&s=09
https://venturebeat.com/ai/nist-staffers-revolt-against-potential-appointment-of-effective-altruist-ai-researcher-to-us-ai-safety-institute/
https://occiglot.github.io/occiglot/posts/occiglot-announcement/
https://twitter.com/EMostaque/status/1767199048337932719?t=tYB3KeabfLlB90XhUX0R7A&s=09
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n [ML News] Devin AI Software Engineer | GPT-4.5-Turbo LEAKED | US Govt Report: Total Extinction](https://i.ytimg.com/vi/q1LrXH5_Oy0/mqdefault.jpg)



![[ML News] GPT-4 Rumors | AI Mind Reading | Neuron Interaction Solved | AI Theorem Proving
#ai #mlnews #gpt4
Your weekly news from the AI & Machine Learning world.
OUTLINE:
0:00 - Introduction
0:25 - AI reads brain signals to predict what youre thinking
3:00 - Closed-form solution for neuron interactions
4:15 - GPT-4 rumors
6:50 - Cerebras supercomputer
7:45 - Meta releases metagenomics atlas
9:15 - AI advances in theorem proving
10:40 - Better diffusion models with expert denoisers
12:00 - BLOOMZ & mT0
13:05 - ICLR reviewers going mad
21:40 - Scaling Transformer inference
22:10 - Infinite nature flythrough generation
23:55 - Blazing fast denoising
24:45 - Large-scale AI training with MultiRay
25:30 - arXiv to include Hugging Face spaces
26:10 - Multilingual Diffusion
26:30 - Music source separation
26:50 - Multilingual CLIP
27:20 - Drug response prediction
27:50 - Helpful Things
ERRATA:
HF did not acquire spaces, they launched spaces themselves and supported Gradio from the start. They later acquired Gradio.
References:
AI reads brain signals to predict what youre thinking
https://mind-vis.github.io/?s=09&utm_source=pocket_saves
https://neurosciencenews.com/bmi-internal-speech-21837/
Closed-form solution for neuron interactions
https://twitter.com/ramin_m_h/status/1592585672606769153/photo/1
https://github.com/raminmh/CfC
https://github.com/raminmh/CfC/blob/main/torch_cfc.py
GPT-4 rumors
https://thealgorithmicbridge.substack.com/p/gpt-4-rumors-from-silicon-valley?utm_source=pocket_reader
Cerebras supercomputer
https://www.cerebras.net/andromeda/
Meta releases metagenomics atlas
https://ai.facebook.com/blog/protein-folding-esmfold-metagenomics/
https://www.genome.gov/genetics-glossary/Metagenomics
AI advances in theorem proving
https://ai.facebook.com/blog/ai-math-theorem-proving/
https://marketplace.visualstudio.com/items?itemName=jroesch.lean
Better diffusion models with expert denoisers
https://deepimagination.cc/eDiffi/
BLOOMZ & mT0
https://arxiv.org/abs/2211.01786?utm_source=pocket_reader
https://huggingface.co/bigscience/bloomz?text=Suggest+at+least+five+related+search+terms+to+%22M%E1%BA%A1ng+neural+nh%C3%A2n+t%E1%BA%A1o%22.
ICLR reviewers going mad
https://twitter.com/XiangruTang/status/1589703605098975237?utm_source=pocket_reader
https://twitter.com/BlancheMinerva/status/1588164585961422849?utm_source=pocket_reader
https://openreview.net/forum?id=pfuqQQCB34
https://twitter.com/peter_richtarik/status/1591408710366408706?utm_source=pocket_reader
Scaling Transformer inference
https://arxiv.org/abs/2211.05102
Infinite nature flythrough generation
https://ai.googleblog.com/2022/11/infinite-nature-generating-3d.html?utm_source=pocket_reader
Blazing fast denoising
https://github.com/dome272/Paella
https://arxiv.org/abs/2211.07292
Large-scale AI training with MultiRay
https://ai.facebook.com/blog/multiray-large-scale-AI-models/
arXiv to include Hugging Face spaces
https://blog.arxiv.org/2022/11/17/discover-state-of-the-art-machine-learning-demos-on-arxiv/
Multilingual Diffusion
https://github.com/FlagAI-Open/FlagAI/tree/master/examples/AltDiffusion
Music source separation
https://github.com/facebookresearch/demucs
https://arxiv.org/abs/2211.08553
Multilingual CLIP
https://twitter.com/rom1504/status/1593719037808320513
Drug response prediction
https://phys.org/news/2022-10-ai-accurately-human-response-drug.html
https://huggingface.co/Onodofthenorth/SD_PixelArt_SpriteSheet_Generator
https://huggingface.co/spaces/ronvolutional/sd-spritesheets
https://github.com/daspartho/prompt-extend
https://huggingface.co/blog/fine-tune-whisper
https://twitter.com/CarsonKatri/status/1585412662724272128
https://github.com/carson-katri/dream-textures/
https://www.youtube.com/playlist?list=PLzvYlJMoZ02Dxtwe-MmH4nOB5jYlMGBjr
https://github.com/xl0/lovely-tensors
https://github.com/jerryjliu/gpt_index
https://colab.research.google.com/drive/1o1qYJcFeywzCIdkfKJy7cTpgZTCM2EI4
https://dagshub.com/blog/launching-data-streaming-and-upload/
https://dagshub.com/blog/build-an-end-2-end-active-learning-pipeline-part-1/
https://github.com/run-ai/genv
https://arxiv.org/abs/2210.14868
https://github.com/timeseriesAI/tsai
https://medium.com/@yangyou_berkeley/diffusion-pretraining-and-hardware-fine-tuning-can-be-almost-7x-cheaper-85e970fe207b
https://medium.com/@hpcaitech/accelerating-structure-prediction-of-protein-monomers-and-multimer-by-11-times-769715dcb5b5
https://github.com/hpcaitech/ColossalAI/tree/main/examples/images/diffusion
https://arxiv.org/abs/2211.03726
https://github.com/Deci-AI/super-gradients
https://github.com/facebookresearch/shumai
https://github.com/huggingface/safetensors
https://github.com/google/learned_optimization/tree/main/learned_optimization/research/general_lopt
https://github.com/NVIDIA-Merlin/dataloader
https://loda-lang.org/
https://loda-lang.org/edit/
https://github.com/EelcoHoogendoorn/numga
https://arxiv.org/abs/2210.07316v1
https://huggingface.co/spaces/mteb/leaderboard
https://twitter.com/natfriedman/status/1575631194032549888
https://github.com/nat/natbot [ML News] GPT-4 Rumors | AI Mind Reading | Neuron Interaction Solved | AI Theorem Proving](https://i.ytimg.com/vi/r8wiBA3ZaQE/mqdefault.jpg)


