A method of transmitting speech signals with reduced bandwith requirements. With this invention an original speech signal is first converted to a textual representation, and a facsimile of the original speech is determined from the textual representation. Then a minimum error turn is derived from the difference between the original speech signal and the facsimile of the original speech signal. The minimum error turn is then compressed, and it is this compressed minimum error turn, along with the textual representation, that is transmitted on the communications medium. At the receiving end, the textual representation and the difference representation are split through a demultiplexer. The textual representation is then passed through a synthesizer while the difference representation is passed through a mapper. The synthesizer along with synthesis parameter storage converts the textual representation into a digital representation of speech, while the mapper modifies the received difference representation by applying sub or super sampling corrections.
A portable communication apparatus allowing increased flexibility and convenience is disclosed. A switch is provided to select one of a voice-character conversion communication mode and a character-voice conversion communication mode depending on a setting instruction. A voice-character converter performs a selected one of a first conversion from voice to character data and a second conversion from character to voice data according to the selected communication mode.
A method of communicating speech across a communication link using very low digital data bandwidth is disclosed, having the steps of: translating speech into text at a source terminal; communicating the text across the communication link to a destination terminal; and translating the text into reproduced speech at the destination terminal. In a preferred embodiment, a speech profile corresponding to the speaker is used to reproduce the speech at the destination terminal so that the reproduced speech more closely approximates the original speech of the speaker. A default voice profile is used to recreate speech when a user profile is unavailable. User specific profiles can be created during training prior to communication or can be created during communication from actual speech. The user profiles can be updated to improve accuracy of recognition and to enhance reproduction of speech. The updated user profiles are transmitted to the destination terminals as needed.
A system and method for IP-based telephone communication utilizing speech-generated text is disclosed. An embodiment of the present invention includes an interface to the Internet for sending and receiving voice and video signals. In addition, the embodiment generates text signals corresponding to voice signals generated by the user. A transmission signal is generated from the text signals and the voice signals for transmission over the Internet. Further, the embodiment of the present invention includes an application which is capable of receiving video and/or speech-generated data transmitted by another device, and concurrently displaying the speech-generated data and video to a user. The speech-generated data may be converted to an audio signal and applied to a speaker concurrently with the video and speech-generated data. In this way, a user is capable of easily communicating speech-generated information to another user during periods of voice signal loss.