When disaster strikes, communication can become a real headache, especially in rural, remote areas where traditional cellular networks and internet services might be out of reach or just plain unreliable. This paper introduces iTantra, an innovative offline Android-based multilingual voice communication system that’s built to ensure reliable communication even over low-bitrate connections. The system works by converting speech into text right on the device, handling local language and message processing, and then packing that information into compact packets for transmission through various communication links. To boost message delivery in these tricky environments, it uses a lightweight reliability mechanism that includes CRC-16 error detection, acknowledgments, retransmission, sequence tracking, and duplicate protection. On the receiving end, packets are reassembled and turned back into speech using offline text-to-speech processing, supporting multiple languages. The prototype operates without needing internet connectivity, showing a speech-to-text latency of about 300–500 ms and an end-to-end latency of roughly 600–1000 ms. In a test involving a 30-second voice message, the compact text representation achieved an impressive 2500:1 reduction in the amount of data transmitted compared to raw audio. This approach highlights the potential for resilient, multilingual, and bandwidth-efficient communication in emergency response situations and other areas where connectivity is limited.institutions.
Introduction
This research proposes iTantra, an open-source, multilingual, offline neural voice transceiver designed for communication in disaster zones, remote rural areas, tactical operations, medical outreach and other environments where cellular or internet connectivity is unavailable. Conventional voice communication requires substantial bandwidth and often depends on cloud services, making it unreliable in disconnected environments. Language diversity further complicates communication among emergency responders, military personnel, healthcare workers and local communities.
The core idea of iTantra is “Transmit Meaning, Not Raw Voice.” Instead of transmitting the complete audio waveform, the system converts speech into compact text using Voice Activity Detection (VAD) and offline Vosk Speech-to-Text (STT). The resulting text is converted into compact binary packets containing language information, priority, sequence numbers and CRC-16 error checking. These packets are transmitted through low-bandwidth channels using a Stop-and-Wait Automatic Repeat reQuest (ARQ) protocol. At the receiver, the text is reconstructed, translated into the user's preferred language using offline Argos Translate or Marian NMT, and converted back to speech using Piper or eSpeak Text-to-Speech (TTS).
The proposed architecture consists of five modular layers: speech acquisition and preprocessing, offline language intelligence, binary packetization, reliable transport/physical abstraction, and receiver-side reassembly, translation and speech synthesis. The transport abstraction allows the same application to operate over Wi-Fi Direct, Bluetooth, VHF/UHF radio modems and USB/serial connections, without changing the higher-level application.
The prototype was developed as a native Android application using Kotlin and Java, with C++ bindings for ONNX Runtime. It operates completely in Airplane Mode, with approximately 280 MB application size and no cloud dependency. The system supports multilingual communication, with the demonstrated prototype providing English, Hindi, Gujarati and Telugu, while the architecture is intended to support additional Indian languages.
Key Results
The experimental results demonstrate substantial communication efficiency:
Parameter
Result
Operating mode
Airplane Mode
Cloud dependency
0%
STT latency
300–500 ms
End-to-end delay
600–1000 ms
Raw 30-sec audio
~500 KB
Transmitted semantic data
~200 bytes
Data reduction
~2500:1
Packet delivery reliability
99%+
Simulated packet loss
30%
Application size
~280 MB
The 2500:1 reduction in transmitted data is the major contribution of the system. Instead of sending approximately 500 KB of audio for a 30-second message, iTantra transmits approximately 200 bytes of textual information. This makes communication much more suitable for low-bitrate links.
The Stop-and-Wait ARQ mechanism, together with CRC-16 error detection, sequence numbers and retransmission, achieved more than 99% delivery reliability even under simulated 30% packet loss. The system also provides an emergency priority mechanism in which a priority 0x03 message can override normal audio settings and play an emergency message at maximum volume.
Main Contribution
The major innovation of iTantra is the combination of semantic compression, offline multilingual NLP and reliable low-bandwidth communication in a single Android platform. Existing consumer applications generally depend on internet connectivity, while many tactical communication systems transmit conventional audio or have limited multilingual capabilities. iTantra instead performs STT, translation and TTS locally, eliminating dependence on cloud infrastructure.
Conclusion
This paper introduces iTantra, a versatile neural voice transceiver that operates independently of infrastructure, works offline, and is adaptable to various transport methods. Its main goal is to tackle the key challenges of limited channel bandwidth and language barriers in remote field settings. By embracing the philosophy of \"Transmit Meaning, Not Raw Voice,\" iTantra transforms spoken audio into compact text right on the device, shrinking a 30-second voice message down to about a 200-byte binary payload. This impressive feat results in a data volume reduction of roughly 2500:1 compared to traditional audio streams.
Tested in full Airplane Mode without relying on the cloud, the system effectively combines on-device Vosk Speech-to-Text (STT), local neural machine translation (NMT), and Piper/eSpeak Text-to-Speech (TTS) engines within a flexible 5-layer architecture. Real-world tests on a functioning two-phone Android prototype show an STT latency of 300–500 ms, an end-to-end delivery latency of 600–1000 ms, and over 99% packet delivery reliability even with a 30% packet loss, thanks to a Stop-and-Wait ARQ protocol supported by 2-byte CRC-16 checksums. Additionally, a 1-byte priority flag guarantees that urgent emergency broadcasts trigger maximum-volume, non-interruptible voice alerts at the receiving end. By facilitating one-to-many broadcast communication with local translation for Indian regional languages, iTantra offers a robust, open-source telecommunications solution for tactical defense, disaster management, and rural healthcare initiatives.
References
[1] eSpeak NG Authors, “eSpeak NG: An open-source speech synthesizer that supports more than hundred languages and accents,” GitHub repository. [Online]. Available: https://github.com/espeak-ng/espeak-ng.
[2] Argos Open Technologies, “Argos Translate: Open-source offline translation library written in Python,” GitHub repository. [Online]. Available: https://github.com/argosopentech/argos-translate.
[3] M. Junczys-Dowmunt, R. Grundkiewicz, T. Dwojak, H. Hoang, K. Heafield, T. Neckermann, F. Seide, U. Germann, A. F. T. Aji, N. Bogoychev, A. F. T. Martins, and A. Birch, “Marian: Fast neural machine translation in C++,” in Proceedings of ACL 2018, System Demonstrations, Melbourne, Australia, 2018, pp. 116–121. doi: 10.18653/v1/P18-4020.
[4] D. Rowe, “Codec 2: Open-source speech codec,” GitHub repository. [Online]. Available: https://github.com/drowe67/codec2.
[5] G. Gerganov, “whisper.cpp: Port of OpenAI’s Whisper model in C/C++,” GitHub repository. [Online]. Available: https://github.com/ggml-org/whisper.cpp.
[6] TAK Product Center, “TAK: Mission success starts with TAK,” U.S. Government. [Online]. Available: https://tak.gov/solutions/military.
[7] C. Schlesiger, Shareable Components Services and Universal Language Translator (ULTRA) Installation on an Android Phone, DEVCOM Army Research Laboratory, 2021. doi: 10.21236/AD1149296.
[8] U.S. Army DEVCOM Army Research Laboratory, “Army researchers partner with Navy to provide expeditionary AI translation tool,” U.S. Army, Mar. 24, 2026. [Online]. Available: https://www.army.mil/article/291315.