Research
Our Research and Publications on AI
Through our research, NetMind seeks to address the most critical challenges and opportunities in AI, and to create a brighter future for humanity by unlocking the potential of AGI.
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
Beyond Closed-Pool Video Retrieval: A Benchmark and Agent Framework for Real-World Video Search and Moment Localization
Context Forcing: Consistent Autoregressive Video Generation with Long Context
VisCoder2: Building Multi-Language Visualization Coding Agents
VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
AceCoder: Acing Coder RL via Automated Test-Case Synthesis
VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
Primitive Vision: Improving Diagram Understanding in MLLMs
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
Enhancing SPARQL Generation by Triplet-order-sensitive Pre-training
A New Dataset and Method for Creativity Assessment Using the Alternate Uses Task
Transform-Equivariant Consistency Learning for Temporal Sentence Grounding
ProS: Facial Omni-Representation Learning via Prototype-Based Self-Distillation
You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding
Filling the Information Gap between Video and Query for Language-Driven Moment Retrieval
Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks
Hypotheses Tree Building for One-Shot Temporal Sentence Localization.
Rethinking the Video Sampling and Reasoning Strategies for Temporal Sentence Grounding.
M³ViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-design
Local Byte Fusion for Neural Machine Translation
RASAT: Integrating Relational Structures into Pretrained Seq2Seq Model for Text-to-SQL.
Backdoor Attacks on Crowd Counting.
Quantifying and alleviating political bias in language models
Unsupervised Temporal Video Grounding with Deep Semantic Clustering.
Memory-Guided Semantic Learning Network for Temporal Sentence Grounding.
Few-shot text classification with triplet networks, data augmentation, and curriculum learning.
Text augmentation in a multi-task view.
Mitigating political bias in language models through reinforced calibration.
An empirical survey of unsupervised text representation methods on twitter data.
What Are People Asking About COVID-19? A Question Classification Dataset
EDA: Easy Data Augmentation techniques for boosting performance on text classification tasks.
Transforming humanity through the power of AI
Partner with us, join the community, or explore opportunities to build the future of AI together.