AI with Papers - Artificial Intelligence & Deep Learning

@ai_deeplearning

Every day fresh updates on Deep Learning, Machine Learning, and Computer Vision (with Papers).

Curated by Alessandro Ferrari | https://www.linkedin.com/in/visionarynet/

Direct contact: @ARGOVISION

AI with Papers - Artificial Intelligence & Deep Learning (English)

Are you passionate about Artificial Intelligence, Deep Learning, and Machine Learning? Look no further! Join the AI with Papers Telegram channel for daily fresh updates on the latest trends and research in the field. Curated by Alessandro Ferrari, a visionary in the world of AI, this channel provides valuable insights and knowledge for enthusiasts and professionals alike. Through curated papers and articles, you will stay informed about the cutting-edge advancements in Computer Vision, Deep Learning, and Machine Learning.

Alessandro Ferrari, the curator of the channel, is a renowned expert in the field of AI and is dedicated to sharing his expertise with the community. With a background in Computer Science and a passion for innovation, he provides valuable content that will keep you at the forefront of the AI revolution.

Whether you are a seasoned professional or a newcomer to the world of AI, the AI with Papers channel has something to offer for everyone. Stay updated on the latest research, trends, and technologies in Artificial Intelligence and Deep Learning. Don't miss out on this incredible opportunity to expand your knowledge and network with like-minded individuals. Join the AI with Papers Telegram channel today and be a part of the future of AI!

Join now and connect with Alessandro Ferrari directly: @ARGOVISION

AI with Papers - Artificial Intelligence & Deep Learning

29 Jan, 07:54

🌅 Generative Human Mesh Recovery 🌅

👉GenHMR is a novel generative framework that reformulates monocular HMR as an image-conditioned generative task, explicitly modeling and mitigating uncertainties in 2D-to-3D mapping process. Impressive results but no code announced 🥺

👉Review https://t.ly/Rrzpj
👉Paper https://arxiv.org/pdf/2412.14444
👉Project m-usamasaleem.github.io/publication/GenHMR/GenHMR.html

1,728

AI with Papers - Artificial Intelligence & Deep Learning

28 Jan, 07:50

☀️ Relightable Full-Body Avatars ☀️

👉#Meta unveils the first approach ever to jointly model the relightable appearance of the body, face, and hands of drivable avatars.

👉Review https://t.ly/kx9gf
👉Paper arxiv.org/pdf/2501.14726
👉Project neuralbodies.github.io/RFGCA

2,192

AI with Papers - Artificial Intelligence & Deep Learning

27 Jan, 13:52

🦕[SOTA] Visual Grounding VOS🦕

👉ReferDINO is the first end-to-end approach for adapting foundational visual grounding models to RVOS. Code & models to be released soon💙

👉Review https://t.ly/SDFy9
👉Paper arxiv.org/pdf/2501.14607
👉Project isee-laboratory.github.io/ReferDINO/
👉Repo github.com/iSEE-Laboratory/ReferDINO

2,329

AI with Papers - Artificial Intelligence & Deep Learning

27 Jan, 09:50

🎨MatAnyone: Human Matting🎨

👉MatAnyone is a novel approach for human video matting that supports the target assignment. Stable tracking in long videos even with complex/ambiguous BGs. Code & 🤗-Demo announced💙

👉Review https://t.ly/NVXsT
👉Paper arxiv.org/pdf/2501.14677
👉Project pq-yang.github.io/projects/MatAnyone
👉Repo TBA

2,639

AI with Papers - Artificial Intelligence & Deep Learning

25 Jan, 12:57

🪆SOTA Points Segmentation🪆

👉VGG Oxford unveils a novel loss to segment objects in videos based on their motion and NO other forms of supervision! Training the net using long-term point trajectories as a supervisory signal to complement optical flow. New SOTA!

👉Review https://t.ly/8Bsbt
👉Paper https://arxiv.org/pdf/2501.12392
👉Code https://github.com/karazijal/lrtl
👉Project www.robots.ox.ac.uk/~vgg/research/lrtl/

3,344

AI with Papers - Artificial Intelligence & Deep Learning

25 Jan, 07:49

🔥 The code of DynOMo is out 🔥

👉DynOMo is a novel model able to track any point in a dynamic scene over time through 3D reconstruction from monocular video: 2D and 3D point tracking from unposed monocular camera input

👉Review https://t.ly/t5pCf
👉Paper https://lnkd.in/dwhzz4_t
👉Repo github.com/dvl-tum/DynOMo
👉Project https://lnkd.in/dMyku2HW

3,511

AI with Papers - Artificial Intelligence & Deep Learning

24 Jan, 14:08

🦠A-Life with Foundation Models🦠

👉A super team unveils ASAL, a new paradigm for Artificial Life research. A diverse range of ALife substrates including Boids, Particle Life, Game of Life, Lenia & Neural Cellular Automata. Code under Apache 2.0💙

👉Review https://t.ly/7SZ8A
👉Paper arxiv.org/pdf/2412.17799
👉Project http://pub.sakana.ai/asal/
👉Repo https://lnkd.in/dP5yxKtw

3,068

AI with Papers - Artificial Intelligence & Deep Learning

24 Jan, 10:24

🎤EMO2: Audio-Driven Avatar🎤

👉Alibaba previews a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Turn your audio ON. Stunning results but no code 🥺

👉Review https://t.ly/x8slQ
👉Paper arxiv.org/pdf/2501.10687
👉Project humanaigc.github.io/emote-portrait-alive-2/
👉Repo 🥺

3,303

AI with Papers - Artificial Intelligence & Deep Learning

23 Jan, 07:37

🧵Time-Aware Pts-Tracking🧵

👉Chrono: feature backbone specifically designed for point tracking with built-in temporal awareness. Long-term temporal context, enabling precise prediction even without the refinements. Code announced💙

👉Review https://t.ly/XAL7G
👉Paper arxiv.orgzpdf/2501.12218
👉Project cvlab-kaist.github.io/Chrono/
👉Repo github.com/cvlab-kaist/Chrono

3,213

AI with Papers - Artificial Intelligence & Deep Learning

22 Jan, 11:38

🔥 [SOTA] Long-Video Depth Anything 🔥

👉ByteDance unveils Video Depth Anything: HQ, consistent depth estimation in SUPER-long videos (over several minutes) without sacrificing efficiency. Based on Depth Anything V2 with a novel efficient spatial-temporal head. Repo available under Apache 2.0💙

👉Review https://t.ly/Q4ZZd
👉Paper arxiv.org/pdf/2501.12375
👉Project https://lnkd.in/dKNwJzbM
👉Repo https://lnkd.in/ddfwwpCj

3,515

AI with Papers - Artificial Intelligence & Deep Learning

22 Jan, 07:01

🌈 #Nvidia Foundation ZS-Stereo 🌈

👉Nvidia unveils FoundationStereo, a foundation model for stereo depth estimation with strong zero-shot generalization. In addition, a large-scale (1M stereo pairs) synthetic training dataset featuring large diversity and high photorealism. Code, model & dataset to be released💙

👉Review https://t.ly/rfBr5
👉Paper arxiv.org/pdf/2501.09898
👉Project nvlabs.github.io/FoundationStereo/
👉Repo github.com/NVlabs/FoundationStereo/tree/master

3,523

AI with Papers - Artificial Intelligence & Deep Learning

21 Jan, 09:26

🧽 Diffusion Video Inpainting 🧽

👉#Alibaba unveils a technical report about DiffuEraser, a video inpainting model based on stable diffusion, designed to fill masked regions with greater details and more coherent structures. Code & weights released under Apache💙

👉Review https://t.ly/7rEll
👉Paper arxiv.org/pdf/2501.10018
👉Project lixiaowen-xw.github.io/DiffuEraser-page/
👉Repo github.com/lixiaowen-xw/DiffuEraser

9,078

AI with Papers - Artificial Intelligence & Deep Learning

20 Jan, 13:19

🏄‍♀️ GSTAR: Gaussian Surface Tracking 🏄‍♀️

👉ETH Zurich unveils GSTAR, a novel framework for photo-realistic rendering, surface reconstruction, and 3D tracking for dynamic scenes while handling topology changes. Code announced💙

👉Review https://t.ly/udpMq
👉Paper arxiv.org/pdf/2501.10283
👉Project chengwei-zheng.github.io/GSTAR/
👉Repo TBA

3,942

AI with Papers - Artificial Intelligence & Deep Learning

18 Jan, 08:03

🎁Free Book: LLM Foundations🎁

👉A fully free book just released on arXiv to outline the basic concepts of #LLMs and related techniques with a focus on the foundational aspects.

✅Chapter 1: basics of pre-training
✅Chapter 2: gen-models & LLMs
✅Chapter 3: prompting methods
✅Chapter 4: alignment methods

👉If you have any background in ML, along with a certain understanding of stuff like Transformers, this book will be "smooth". However, even without this prior knowledge, it is still perfectly fine because the contents of each chapter are self-contained.

👉Review https://t.ly/9LGCa
👉Book https://lnkd.in/d3VkswZf

4,947

AI with Papers - Artificial Intelligence & Deep Learning

17 Jan, 13:09

🔥 GAGA: Group Any Gaussians 🔥

👉GAGA is a framework that reconstructs and segments open-world 3D scenes by leveraging inconsistent 2D masks predicted by zero-shot segmentation models. Code available, recently updated💙

👉Review https://t.ly/Nk_jT
👉Paper www.gaga.gallery/static/pdf/Gaga.pdf
👉Project www.gaga.gallery/
👉Repo github.com/weijielyu/Gaga

4,939

AI with Papers - Artificial Intelligence & Deep Learning

16 Jan, 07:55

🧞‍♂️Omni-RGPT: SOTA MLLM Understanding🧞‍♂️

👉 #NVIDIA presents Omni-RGPT, MLLM for region-level comprehension for both images & videos. New SOTA on image/video-based commonsense reasoning.

👉Review https://t.ly/KHnQ7
👉Paper arxiv.org/pdf/2501.08326
👉Project miranheo.github.io/omni-rgpt/
👉Repo TBA soon

5,076

AI with Papers - Artificial Intelligence & Deep Learning

15 Jan, 14:58

🆘 Help: Looking for Outstanding Speakers 🆘

👉Who would you suggest as a speaker for your ideal conference on AI (CV, LLM, RAG, ML, HW Optimization, AI & Space, etc.)? Only “hardcore” technical talks, no commercial at all. Please comment here with name, topic and affiliation (es: Paul Gascoigne, Computer Vision & Football, Scotland Team).

⭐Guaranteed tickets & more for the suggestions that will become invited speakers ;)

4,485

AI with Papers - Artificial Intelligence & Deep Learning

14 Jan, 12:42

🏆Universal Detector-Free Match🏆

👉MatchAnything: novel detector-free universal matcher across unseen real-world single/cross-modality domains. Same weights for everything. Code announced, to be released 💙

👉Review https://t.ly/sx92L
👉Paper https://lnkd.in/dWwRwGyY
👉Project https://lnkd.in/dCwb2Yte
👉Repo https://lnkd.in/dnUXYzQ5

5,422

AI with Papers - Artificial Intelligence & Deep Learning

14 Jan, 09:02

❤️‍🔥 Uncommon object in #3D ❤️‍🔥

👉#META releases uCO3D, a new object-centric dataset for 3D AI. The largest publicly-available collection of HD videos of objects with 3D annotations that ensures full-360◦ coverage. Code & data under CCA 4.0💙

👉Review https://t.ly/Z_tvA
👉Paper https://arxiv.org/pdf/2501.07574
👉Project https://uco3d.github.io/
👉Repo github.com/facebookresearch/uco3d

4,548

AI with Papers - Artificial Intelligence & Deep Learning

10 Jan, 07:33

🔥 Depth Any Camera (SOTA) 🔥

👉DAC is a novel and powerful zero-shot metric depth estimation framework that extends a perspective-trained model to effectively handle cams with varying FoVs (including large fisheye & 360◦). Code announced (not available yet)💙

👉Review https://t.ly/1qz4F
👉Paper arxiv.org/pdf/2501.02464
👉Project yuliangguo.github.io/depth-any-camera/
👉Repo github.com/yuliangguo/depth_any_camera

4,561

AI with Papers - Artificial Intelligence & Deep Learning

09 Jan, 07:40

⚽ FIFA 3D Human Pose ⚽

👉#FIFA WorldPose is a novel dataset for multi-person global pose estimation in the wild, featuring footage from the 2022 World Cup. 2.5M+ annotation, released 💙

👉Review https://t.ly/kvGVQ
👉Paper arxiv.org/pdf/2501.02771
👉Project https://lnkd.in/d5hFWpY2
👉Dataset https://lnkd.in/dAphJ9WA

4,233

AI with Papers - Artificial Intelligence & Deep Learning

08 Jan, 09:18

🧤World-Space Ego 3D Hands🧤

👉The Imperial College unveils HaWoR, a novel world-space 3D hand motion estimation for egocentric videos. The new SOTA on both cam pose estimation & hand motion reconstruction. Code under Attribution-NC-ND 4.0 Int.💙

👉Review https://t.ly/ozJn7
👉Paper arxiv.org/pdf/2501.02973
👉Project hawor-project.github.io/
👉Code github.com/ThunderVVV/HaWoR

3,733

AI with Papers - Artificial Intelligence & Deep Learning

07 Jan, 09:12

🥮 SOTA probabilistic tracking🥮

👉ProTracker is a novel framework for robust and accurate long-term dense tracking of arbitrary points in videos. Code released under CC Attribution-NonCommercial💙

👉Review https://t.ly/YY_PH
👉Paper https://arxiv.org/pdf/2501.03220
👉Project michaelszj.github.io/protracker/
👉Code github.com/Michaelszj/pro-tracker

3,744

AI with Papers - Artificial Intelligence & Deep Learning

05 Jan, 13:46

AI with Papers - Artificial Intelligence & Deep Learning pinned «What is your favorite source for the AI updates?»

AI with Papers - Artificial Intelligence & Deep Learning

05 Jan, 12:09

⭐ Poll Alert!! ⭐

[EDIT] see below

4,570

AI with Papers - Artificial Intelligence & Deep Learning

04 Jan, 10:13

🌳 HD Video Object Insertion 🌳

👉VideoAnydoor is a novel zero-shot video object insertion #AI with high-fidelity detail preservation and precise motion control. All-in-one: video VTON, face swapping, logo insertion, multi-region editing, etc.

👉Review https://t.ly/hyvRq
👉Paper arxiv.org/pdf/2501.01427
👉Project videoanydoor.github.io/
👉Repo TBA

5,093

AI with Papers - Artificial Intelligence & Deep Learning

31 Dec, 07:57

⭐TOP 10 Papers you loved - 2024⭐

👉Here the list of my posts you liked the most in 2024, thank you all 💙

𝐏𝐚𝐩𝐞𝐫𝐬:
⭐"Look Ma, no markers"
⭐T-Rex 2 Detector
⭐Models at Any Resolution

👉The full list with links: https://t.ly/GvQVy

6,061

AI with Papers - Artificial Intelligence & Deep Learning

28 Dec, 10:20

🔄️ Orient Anything in 3D 🔄️
️
👉Orient Anything is a novel robust image-based object orientation estimation model. By training on 2M rendered labeled images, it achieves strong zero-shot generalization in the wild. Code released💙

👉Review https://t.ly/ro5ep
👉Paper arxiv.org/pdf/2412.18605
👉Project orient-anything.github.io/
👉Code https://lnkd.in/d_3k6Nxz

6,318

AI with Papers - Artificial Intelligence & Deep Learning

23 Dec, 09:11

🍄 Open-MLLMs Self-Driving 🍄

👉OpenEMMA: a novel open-source e2e framework based on MLLMs (via Chain-of-Thought reasoning). Effectiveness, generalizability, and robustness across a variety of challenging driving scenarios. Code released under Apache 2.0💙

👉Review https://t.ly/waLZI
👉Paper https://arxiv.org/pdf/2412.15208
👉Code https://github.com/taco-group/OpenEMMA

7,181

AI with Papers - Artificial Intelligence & Deep Learning

19 Dec, 12:55

🫶 Dynamic Cam-4D Hands 🫶

👉The Imperial College unveils Dyn-HaMR, the first approach to reconstruct 4D global hand motion from monocular videos recorded by dynamic cameras in the wild. Code announced under MIT💙

👉Review https://t.ly/h5vV7
👉Paper arxiv.org/pdf/2412.12861
👉Project dyn-hamr.github.io/
👉Repo github.com/ZhengdiYu/Dyn-HaMR

7,846

AI with Papers - Artificial Intelligence & Deep Learning

13 Dec, 14:02

🐕 Gaze-LLE: Neural Gaze 🐕

👉Gaze-LLE: novel transformer framework that streamlines gaze target by leveraging features from frozen DINOv2 encoder. Code & models under MIT 💙

👉Review https://t.ly/SadoF
👉Paper arxiv.org/pdf/2412.09586
👉Repo github.com/fkryan/gazelle

8,459

AI with Papers - Artificial Intelligence & Deep Learning

12 Dec, 14:38

🌹 4D Neural Templates 🌹

👉#Stanford unveils Neural Templates, generating HQ temporal object intrinsics for several natural phenomena and enable the sampling and controllable rendering of these dynamic objects from any viewpoint, at any time of their lifespan. A novel task in vision is born💙

👉Review https://t.ly/ka_Qf
👉Paper https://arxiv.org/pdf/2412.05278
👉Project https://chen-geng.com/rose4d#toi

7,076

AI with Papers - Artificial Intelligence & Deep Learning

11 Dec, 12:56

🦢 Track4Gen: Diffusion + Tracking 🦢

👉Track4Gen: spatially aware video generator that combines video diffusion loss with point tracking across frames, providing enhanced spatial supervision on the diffusion features. GenAI with points-based motion control. Stunning results but no code announced😢

👉Review https://t.ly/9ujhc
👉Paper arxiv.org/pdf/2412.06016
👉Project hyeonho99.github.io/track4gen/
👉Gallery hyeonho99.github.io/track4gen/full.html

6,495

AI with Papers - Artificial Intelligence & Deep Learning

10 Dec, 07:50

🧤GigaHands: Massive #3D Hands🧤

👉Novel massive #3D bimanual activities dataset: 34 hours of activities, 14k hand motions clips paired with 84k text annotation, 183M+ unique hand images

👉Review https://t.ly/SA0HG
👉Paper www.arxiv.org/pdf/2412.04244
👉Repo github.com/brown-ivl/gigahands
👉Project ivl.cs.brown.edu/research/gigahands.html

5,892

AI with Papers - Artificial Intelligence & Deep Learning

05 Dec, 08:07

🦘AniGS: Single Pic Animatable Avatar🦘

👉#Alibaba unveils AniGS: given a single human image as input it rebuilds a Hi-Fi 3D avatar in a canonical pose, which can be used for both photorealistic rendering & real-time animation. Source code announced, to be released💙

👉Review https://t.ly/4yfzn
👉Paper arxiv.org/pdf/2412.02684
👉Project lingtengqiu.github.io/2024/AniGS/
👉Repo github.com/aigc3d/AniGS

6,784

AI with Papers - Artificial Intelligence & Deep Learning

04 Dec, 08:12

🌈Motion Prompting Video Generation🌈

👉DeepMind unveils ControlNet, novel video generation model conditioned on spatio-temporally sparse or dense motion trajectories. Amazing results, but no code announced 😢

👉Review https://t.ly/VyKbv
👉Paper arxiv.org/pdf/2412.02700
👉Project motion-prompting.github.io

6,257

AI with Papers - Artificial Intelligence & Deep Learning

03 Dec, 07:46

⚽Universal Soccer Foundation Model⚽

👉Universal Soccer Video Understanding: SoccerReplay-1988 - the largest multi-modal soccer dataset - and MatchVision - the first vision-lang. foundation models for soccer. Code, dataset & checkpoints to be released💙

👉Review https://t.ly/-X90B
👉Paper https://arxiv.org/pdf/2412.01820
👉Project https://jyrao.github.io/UniSoccer/
👉Repo https://github.com/jyrao/UniSoccer

6,172

AI with Papers - Artificial Intelligence & Deep Learning

02 Dec, 07:58

🔥Video Depth without Video Models🔥

👉RollingDepth: turning a single-image latent diffusion model (LDM) into the novel SOTA depth estimator. It works better than dedicated model for depth 🤯 Code under Apache💙

👉Review https://t.ly/R4LqS
👉Paper https://arxiv.org/pdf/2411.19189
👉Project https://rollingdepth.github.io/
👉Repo https://github.com/prs-eth/rollingdepth

5,805

AI with Papers - Artificial Intelligence & Deep Learning

29 Nov, 09:12

👺HiFiVFS: Extreme Face Swapping👺

👉HiFiVFS: HQ face swapping videos even in extremely challenging scenarios (occlusion, makeup, lights, extreme poses, etc.). Impressive results, no code announced😢

👉Review https://t.ly/ea8dU
👉Paper https://arxiv.org/pdf/2411.18293
👉Project https://cxcx1996.github.io/HiFiVFS

6,752

AI with Papers - Artificial Intelligence & Deep Learning

28 Nov, 13:26

🧶SOTA track-by-propagation🧶

👉SambaMOTR is a novel e2e model (based on Samba) for long-range dependencies and interactions between tracklets to handle complex motion patterns / occlusions. Code in Jan. 25 💙

👉Review https://t.ly/QSQ8L
👉Paper arxiv.org/pdf/2410.01806
👉Project sambamotr.github.io/
👉Repo https://lnkd.in/dRDX6nk2

6,409

AI with Papers - Artificial Intelligence & Deep Learning

27 Nov, 09:25

🛟 StableAnimator: ID-aware Humans 🛟

👉StableAnimator: first e2e ID-preserving diffusion for HQ videos without any post-processing. Input: single image + sequence of poses. Insane results!

👉Review https://t.ly/JDtL3
👉Paper https://arxiv.org/pdf/2411.17697
👉Project francis-rings.github.io/StableAnimator/
👉Code github.com/Francis-Rings/StableAnimator

6,197

AI with Papers - Artificial Intelligence & Deep Learning

26 Nov, 13:55

🦙 EdgeCape: SOTA Agnostic Pose 🦙

👉EdgeCap: new SOTA in Category-Agnostic Pose Estimation (CAPE): finding keypoints across diverse object categories using only one or a few annotated support images. Source code released💙

👉Review https://t.ly/4TpAs
👉Paper https://arxiv.org/pdf/2411.16665
👉Project https://orhir.github.io/edge_cape/
👉Code https://github.com/orhir/EdgeCape

5,853

AI with Papers - Artificial Intelligence & Deep Learning

26 Nov, 08:27

🌎All Languages Matter: LMMs vs. 100 Lang.🌎

👉ALM-Bench aims to assess the next generation of massively multilingual multimodal models in a standardized way, pushing the boundaries of LMMs towards better cultural understanding and inclusivity. Code & Dataset 💙

👉Review https://t.ly/VsoJB
👉Paper https://lnkd.in/ddVVZfi2
👉Project https://lnkd.in/dpssaeRq
👉Code https://lnkd.in/dnbaJJE4
👉Dataset https://lnkd.in/drw-_95v

5,341

AI with Papers - Artificial Intelligence & Deep Learning

23 Nov, 09:09

🦖Dino-X: Unified Obj-Centric LVM🦖

👉Unified vision model for Open-World Detection, Segmentation, Phrase Grounding, Visual Counting, Pose, Prompt-Free Detection/Recognition, Dense Caption, & more. Demo & API announced 💙

👉Review https://t.ly/CSQon
👉Paper https://lnkd.in/dc44ZM8v
👉Project https://lnkd.in/dehKJVvC
👉Repo https://lnkd.in/df8Kb6iz

6,404

AI with Papers - Artificial Intelligence & Deep Learning

22 Nov, 07:40

⚔️SAMurai: SAM for Tracking⚔️

👉UWA unveils SAMURAI, an enhanced adaptation of SAM 2 specifically designed for visual object tracking. New SOTA! Code under Apache 2.0💙

👉Review https://t.ly/yGU0P
👉Paper https://arxiv.org/pdf/2411.11922
👉Repo https://github.com/yangchris11/samurai
👉Project https://yangchris11.github.io/samurai/

6,635

AI with Papers - Artificial Intelligence & Deep Learning

18 Nov, 10:31

🧰 EchoMimicV2: Semi-body Human 🧰

👉Alipay (ANT Group) unveils EchoMimicV2, the novel SOTA half-body human animation via APD-Harmonization. See clip with audio (ZH/ENG). Code & Demo announced💙

👉Review https://t.ly/enLxJ
👉Paper arxiv.org/pdf/2411.10061
👉Project antgroup.github.io/ai/echomimic_v2/
👉Repo-v2 github.com/antgroup/echomimic_v2
👉Repo-v1 https://github.com/antgroup/echomimic

7,155

AI with Papers - Artificial Intelligence & Deep Learning

15 Nov, 13:47

🧶 MagicQuill: super-easy Diffusion Editing 🧶

👉MagicQuill is a novel system designed to support users in smart editing of images. Robust UI/UX (e.g., inserting/erasing objects, colors, etc.) under a multimodal LLM to anticipate user intentions in real time. Code & Demos released 💙

👉Review https://t.ly/hJyLa
👉Paper https://arxiv.org/pdf/2411.09703
👉Project https://magicquill.art/demo/
👉Repo https://github.com/magic-quill/magicquill
👉Demo https://huggingface.co/spaces/AI4Editing/MagicQuill

6,969

AI with Papers - Artificial Intelligence & Deep Learning

15 Nov, 07:32

🛥️ Global Tracklet Association MOT 🛥️

👉A novel universal, model-agnostic method designed to refine and enhance tracklet association for single-camera MOT. Suitable for datasets such as SportsMOT, SoccerNet & similar. Source code released💙

👉Review https://t.ly/gk-yh
👉Paper https://lnkd.in/dvXQVKFw
👉Repo https://lnkd.in/dEJqiyWs

6,490

AI with Papers - Artificial Intelligence & Deep Learning

14 Nov, 07:54

🔥 4 NanoSeconds inference 🔥

👉LogicTreeNet: convolutional differentiable logic gate net. with logic gate tree kernels: Computer Vision into differentiable LGNs. Up to 6100% smaller than SOTA, inference in 4 NANOsecs!

👉Review https://t.ly/GflOW
👉Paper https://lnkd.in/dAZQr3dW
👉Full clip https://lnkd.in/dvDJ3j-u

4,866

AI with Papers - Artificial Intelligence & Deep Learning

13 Nov, 07:51

🐔SeedEdit: foundational T2I🐔

👉ByteDance unveils a novel T2I foundational model capable of delivering stable, high-aesthetic image edits which maintain image quality through unlimited rounds of editing instructions. No code announced but a Demo is online💙

👉Review https://t.ly/hPlnN
👉Paper https://arxiv.org/pdf/2411.06686
👉Project team.doubao.com/en/special/seededit
🤗Demo https://huggingface.co/spaces/ByteDance/SeedEdit-APP

4,342

AI with Papers - Artificial Intelligence & Deep Learning

11 Nov, 13:47

❄️Don’t Look Twice: ViT by RLT❄️

👉CMU unveils RLT: speeding up the video transformers inspired by run-length encoding for data compression. Speed the training up and reducing the token count by up to 80%! Source Code announced 💙

👉Review https://t.ly/ccSwN
👉Paper https://lnkd.in/d6VXur_q
👉Project https://lnkd.in/d4tXwM5T
👉Repo TBA

5,074

AI with Papers - Artificial Intelligence & Deep Learning

10 Nov, 10:43

🫠 X-Portrait 2: SOTA(?) Portrait Animation 🫠

👉ByteDance unveils a preview of X-Portrait2, the new SOTA expression encoder model that implicitly encodes every minuscule expressions from the input by training it on large-scale datasets. Impressive results but no paper & code announced.

👉Review https://t.ly/8Owh9 [UPDATE]
👉Paper ?
👉Project byteaigc.github.io/X-Portrait2/
👉Repo ?

5,121

AI with Papers - Artificial Intelligence & Deep Learning

08 Nov, 09:30

🧠 Single Neuron Reconstruction 🧠

👉SIAT unveils NeuroFly, a framework for large-scale single neuron reconstruction. Formulating neuron reconstruction task as a 3-stage streamlined workflow: automatic segmentation - connection - manual proofreading. Bridging computer vision and neuroscience 💙

👉Review https://t.ly/Y5Xu0
👉Paper https://arxiv.org/pdf/2411.04715
👉Repo github.com/beanli161514/neurofly

5,028

AI with Papers - Artificial Intelligence & Deep Learning

07 Nov, 08:24

💪 Muscles in Time Dataset 💪

👉Muscles in Time (MinT) is a large-scale synthetic muscle activation dataset. MinT contains 9+ hours of simulation data covering 227 subjects and 402 simulated muscle strands. Code & Dataset available soon 💙

👉Review https://t.ly/108g6
👉Paper arxiv.org/pdf/2411.00128
👉Project davidschneider.ai/mint
👉Code github.com/simplexsigil/MusclesInTime

5,031

AI with Papers - Artificial Intelligence & Deep Learning

05 Nov, 07:22

🏣 CityGaussianV2: Large-Scale City 🏣

👉A novel approach for large-scale scene reconstruction that addresses critical challenges related to geometric accuracy and efficiency: 10x compression, 25% faster & -50% memory! Source code released💙

👉Review https://t.ly/Xgn59
👉Paper arxiv.org/pdf/2411.00771
👉Project dekuliutesla.github.io/CityGaussianV2/
👉Code github.com/DekuLiuTesla/CityGaussian

5,904

AI with Papers - Artificial Intelligence & Deep Learning

04 Nov, 07:39

☀️ Universal Relightable Avatars ☀️

👉#Meta unveils URAvatar, photorealistic & relightable avatars from phone scan with unknown illumination. Stunning results!

👉Review https://t.ly/U-ESX
👉Paper arxiv.org/pdf/2410.24223
👉Project junxuan-li.github.io/urgca-website

5,576

AI with Papers - Artificial Intelligence & Deep Learning

01 Nov, 07:30

🍜 REM: Segment What You Describe 🍜

👉REM is a framework for segmenting concepts in video that can be described via LLM. Suitable for rare & non-object dynamic concepts, such as waves, smoke, etc. Code & Data announced 💙

👉Review https://t.ly/OyVtV
👉Paper arxiv.org/pdf/2410.23287
👉Project https://miccooper9.github.io/projects/ReferEverything/

7,150

AI with Papers - Artificial Intelligence & Deep Learning

31 Oct, 08:13

🔥🔥 The code is out 🔥🔥

👉Code https://github.com/HaixinShi/fmov_pose

6,391

AI with Papers - Artificial Intelligence & Deep Learning

31 Oct, 08:00

🔥 D-FINE: new SOTA Detector 🔥

👉D-FINE, a powerful real-time object detector that achieves outstanding localization precision by redefining the bounding box regression task in DETR model. New SOTA on MS COCO with additional data. Code & models available 💙

👉Review https://t.ly/aw9fN
👉Paper https://arxiv.org/pdf/2410.13842
👉Code https://github.com/Peterande/D-FINE

5,267

AI with Papers - Artificial Intelligence & Deep Learning

29 Oct, 07:53

🫐 Blendify: #Python + Blender 🫐

👉Lightweight Python framework that provides a high-level API for creating & rendering scenes with #Blender. It simplifies data augmentation & synthesis. Source Code released💙

👉Review https://t.ly/l0crA
👉Paper https://arxiv.org/pdf/2410.17858
👉Code https://virtualhumans.mpi-inf.mpg.de/blendify/

5,240

AI with Papers - Artificial Intelligence & Deep Learning

25 Oct, 10:49

⛈️ SMITE: SEGMENT IN TIME ⛈️

👉SFU unveils SMITE: a novel AI that -with only one or few segmentation references with fine granularity- is able to segment different unseen videos respecting the segmentation references. Dataset & Code (under Apache 2.0) announced 💙

👉Review https://t.ly/w6aWJ
👉Paper arxiv.org/pdf/2410.18538
👉Project segment-me-in-time.github.io/
👉Repo github.com/alimohammadiamirhossein/smite

5,974

AI with Papers - Artificial Intelligence & Deep Learning

24 Oct, 09:05

🌻 Plant Camouflage Detection🌻

👉PlantCamo Dataset is the first dataset for plant camouflage detection: 1,250 images with camouflage characteristics. Source Code released 💙

👉Review https://t.ly/pYFX4
👉Paper arxiv.org/pdf/2410.17598
👉Code github.com/yjybuaa/PlantCamo

6,396

AI with Papers - Artificial Intelligence & Deep Learning

23 Oct, 06:48

🪁 PL2Map: efficient neural 2D-3D 🪁

👉PL2Map is a novel neural network tailored for efficient representation of complex point & line maps. A natural representation of 2D-3D correspondences

👉Review https://t.ly/D-bVD
👉Paper arxiv.org/pdf/2402.18011
👉Project https://thpjp.github.io/pl2map
👉Code https://github.com/ais-lab/pl2map

5,989

AI with Papers - Artificial Intelligence & Deep Learning

19 Oct, 07:08

🧿 Look Ma, no markers 🧿

👉#Microsoft unveils the first technique for marker-free, HQ reconstruction of COMPLETE human body, including eyes and tongue, without requiring any calibration, manual intervention or custom hardware. Impressive results! Repo for training & Dataset released💙

👉Review https://t.ly/5fN0g
👉Paper arxiv.org/pdf/2410.11520
👉Project microsoft.github.io/SynthMoCap/
👉Repo github.com/microsoft/SynthMoCap

6,533

AI with Papers - Artificial Intelligence & Deep Learning

18 Oct, 12:30

🔥BitNet: code of 1-bit LLM released🔥

👉BitNet by #Microsoft, announced in late 2023, is a 1-bit Transformer architecture designed for LLMs. BitLinear as a drop-in replacement of the nn.Linear layer in order to train 1-bit weights from scratch. Source Code just released 💙

👉Review https://t.ly/3G2LA
👉Paper arxiv.org/pdf/2310.11453
👉Code https://lnkd.in/duPADJVb

6,184

AI with Papers - Artificial Intelligence & Deep Learning

18 Oct, 07:10

☀️ GS + Depth = SOTA ☀️

👉DepthSplat, the new SOTA in depth estimation & novel view synthesis. The key feature is the cross-task interaction between Gaussian Splatting & depth estimation. Source Code to be released soon💙

👉Review https://t.ly/87HuH
👉Paper arxiv.org/abs/2410.13862
👉Project haofeixu.github.io/depthsplat/
👉Code github.com/cvg/depthsplat

5,371

AI with Papers - Artificial Intelligence & Deep Learning

17 Oct, 06:56

🦠 Neural Metamorphosis 🦠

👉NU Singapore unveils NeuMeta to transform neural nets by allowing a single model to adapt on the fly to different sizes, generating the right weights when needed.

👉Review https://t.ly/DJab3
👉Paper arxiv.org/pdf/2410.11878
👉Project adamdad.github.io/neumeta
👉Code github.com/Adamdad/neumeta

5,368

AI with Papers - Artificial Intelligence & Deep Learning

16 Oct, 08:43

🔥 CoTracker3 by #META is out! 🔥

👉#Meta (+VGG Oxford) unveils CoTracker3, a new tracker that outperforms the previous SoTA by a large margin using only the 0.1% of the training data 🤯🤯🤯

👉Review https://t.ly/TcRIv
👉Paper arxiv.org/pdf/2410.11831
👉Project cotracker3.github.io/
👉Code github.com/facebookresearch/co-tracker

5,264

AI with Papers - Artificial Intelligence & Deep Learning

16 Oct, 07:26

🪞Robo-Emulation via Video Imitation🪞

👉OKAMI (UT & #Nvidia) is a novel foundation method that generates a manipulation plan from a single RGB-D video and derives a policy for execution.

👉Review https://t.ly/_N29-
👉Paper arxiv.org/pdf/2410.11792
👉Project https://lnkd.in/d6bHF_-s

4,591

AI with Papers - Artificial Intelligence & Deep Learning

15 Oct, 06:50

🔥 DEPTH ANY VIDEO is out! 🔥

👉DAV is a novel foundation model for image/video depth estimation.The new SOTA for accuracy & consistency, up to 150 FPS!

👉Review https://t.ly/CjSz2
👉Paper arxiv.org/pdf/2410.10815
👉Project depthanyvideo.github.io/
👉Code github.com/Nightmare-n/DepthAnyVideo

5,444

AI with Papers - Artificial Intelligence & Deep Learning

14 Oct, 13:25

🥎POKEFLEX: Soft Object Dataset🥎

👉PokeFlex from ETH is a dataset that includes 3D textured meshes, point clouds, RGB & depth maps of deformable objects. Pretrained models & dataset announced💙

👉Review https://t.ly/GXggP
👉Paper arxiv.org/pdf/2410.07688
👉Project https://lnkd.in/duv-jS7a
👉Repo

5,058

AI with Papers - Artificial Intelligence & Deep Learning

12 Oct, 06:09

💡Diffusion Models Relighting💡

👉#Netflix unveils DifFRelight, a novel free-viewpoint facial relighting via diffusion model. Precise lighting control, high-fidelity relit facial images from flat-lit inputs.

👉Review https://t.ly/fliXU
👉Paper arxiv.org/pdf/2410.08188
👉Project www.eyelinestudios.com/research/diffrelight.html

5,051

AI with Papers - Artificial Intelligence & Deep Learning

10 Oct, 05:27

🥦Gaussian Splatting VTON🥦

👉GS-VTON is a novel image-prompted 3D-VTON which, by leveraging 3DGS as the 3D representation, enables the transfer of pre-trained knowledge from 2D VTON models to 3D while improving cross-view consistency. Code announced💙

👉Review https://t.ly/sTPbW
👉Paper arxiv.org/pdf/2410.05259
👉Project yukangcao.github.io/GS-VTON/
👉Repo github.com/yukangcao/GS-VTON

5,586

AI with Papers - Artificial Intelligence & Deep Learning

09 Oct, 06:28

🐏 EFM3D: 3D Ego-Foundation 🐏

👉#META presents EFM3D, the first benchmark for 3D object detection and surface regression on HQ annotated egocentric data of Project Aria. Datasets & Code released💙

👉Review https://t.ly/cDJv6
👉Paper arxiv.org/pdf/2406.10224
👉Project www.projectaria.com/datasets/aeo/
👉Repo github.com/facebookresearch/efm3d

5,213

AI with Papers - Artificial Intelligence & Deep Learning

04 Oct, 06:51

🔥 "Deep Gen-AI" Full Course 🔥

👉A fresh course from Stanford about the probabilistic foundations and algorithms for deep generative models. A novel overview about the evolution of the genAI in #computervision, language and more...

👉Review https://t.ly/ylBxq
👉Course https://lnkd.in/dMKH9gNe
👉Lectures https://lnkd.in/d_uwDvT6

6,024

AI with Papers - Artificial Intelligence & Deep Learning

03 Oct, 14:34

🛳️ EVER Ellipsoid Rendering 🛳️

👉UCSD & Google present EVER, a novel method for real-time differentiable emission-only volume rendering. Unlike 3DGS it does not suffer from popping artifacts and view dependent density, achieving ∼30 FPS at 720p on #NVIDIA RTX4090.

👉Review https://t.ly/zAfGU
👉Paper arxiv.org/pdf/2410.01804
👉Project half-potato.gitlab.io/posts/ever/

6,055

AI with Papers - Artificial Intelligence & Deep Learning

03 Oct, 09:59

🦴 One-Image Object Detection 🦴

👉Delft University (+Hensoldt Optronics) introduces OSSA, a novel unsupervised domain adaptation method for object detection that utilizes a single, unlabeled target image to approximate the target domain style. Code released💙

👉Review https://t.ly/-li2G
👉Paper arxiv.org/pdf/2410.00900
👉Code github.com/RobinGerster7/OSSA

5,767

AI with Papers - Artificial Intelligence & Deep Learning

01 Oct, 12:48

🍇SPARK: Real-time Face Capture🍇

👉Technicolor Group unveils SPARK, a novel high-precision 3D face capture via collection of unconstrained videos of a subject as prior information. New SOTA able to handle unseen pose, expression and lighting. Impressive results. Code & Model announced💙

👉Review https://t.ly/rZOgp
👉Paper arxiv.org/pdf/2409.07984
👉Project kelianb.github.io/SPARK/
👉Repo github.com/KelianB/SPARK/

5,752

AI with Papers - Artificial Intelligence & Deep Learning

30 Sep, 11:57

👩‍🦰 SOTA Gaussian Haircut 👩‍🦰

👉ETH et. al unveils Gaussian Haircut, the new SOTA in hair reconstruction via dual representation (classic + 3D Gaussian). Code and Model announced💙

👉Review https://t.ly/aiOjq
👉Paper arxiv.org/pdf/2409.14778
👉Project https://lnkd.in/dFRm2ycb
👉Repo https://lnkd.in/d5NWNkb5

5,275

AI with Papers - Artificial Intelligence & Deep Learning

24 Sep, 13:42

🌾 New SOTA Edge Detection 🌾

👉CUP (+ ESPOCH) unveils the new SOTA for Edge Detection (NBED); superior performance consistently across multiple benchmarks, even compared with huge computational cost and complex training models. Source Code released💙

👉Review https://t.ly/zUMcS
👉Paper arxiv.org/pdf/2409.14976
👉Code github.com/Li-yachuan/NBED

6,549

AI with Papers - Artificial Intelligence & Deep Learning

24 Sep, 08:13

🩰 Dressed Humans in the wild 🩰

👉ETH (+ #Microsoft ) ReLoo: novel 3D-HQ reconstruction of humans dressed in loose garments from mono in-the-wild clips. No prior assumptions about the garments. Source Code announced, coming 💙

👉Review https://t.ly/evgmN
👉Paper arxiv.org/pdf/2409.15269
👉Project moygcc.github.io/ReLoo/
👉Code github.com/eth-ait/ReLoo

5,894

AI with Papers - Artificial Intelligence & Deep Learning

23 Sep, 13:40

🎢 Robo-quadruped Parkour🎢

👉LAAS-CNRS unveils a novel RL approach to perform agile skills that are reminiscent of parkour, such as walking, climbing high steps, leaping over gaps, and crawling under obstacles. Data and Code available💙

👉Review https://t.ly/-6VRm
👉Paper arxiv.org/pdf/2409.13678
👉Project gepetto.github.io/SoloParkour/
👉Code github.com/Gepetto/SoloParkour

5,324

AI with Papers - Artificial Intelligence & Deep Learning

23 Sep, 12:46

🌏 JoyHallo: Mandarin Digital Human 🌏

👉JD Health faced the challenges of audio-driven video generation in Mandarin, a task complicated by the language’s intricate lip movements and the scarcity of HQ datasets. Impressive results (-> audio ON). Code Models available💙

👉Review https://t.ly/5NGDh
👉Paper arxiv.org/pdf/2409.13268
👉Project jdh-algo.github.io/JoyHallo/
👉Code github.com/jdh-algo/JoyHallo

4,922