1. A method to be executed at least in part in a computing device for performing concatenative speech synthesis, the method comprising:
determining feature vectors for speech segments based on a matrix of concatenation costs;
applying distance weighting to each speech segment pair based on the feature vectors;
clustering the speech segments into a predefined number of groups such that an average distance between speech segments within each group is minimized;
selecting a representative speech segment for each group; and
generating a compressed concatenation cost matrix based on the representative speech segments.
2. The method of claim 1, further comprising:
pre-saving the compressed concatenation cost matrix for real time computations in synthesizing speech.
3. The method of claim 1, wherein the distance weighting is applied employing one of: a Euclidean distance function and a city block distance function.
4. The method of claim 1, wherein the matrix of concatenation costs is constructed along a preceding speech segments axis and a following speech segments axis.
5. The method of claim 4, wherein a concatenation cost between a preceding speech segment and a following speech segment is different from a concatenation cost between the same speech segments with an order of the speech segments reversed.
6. The method of claim 1, wherein the representative speech segment for each group is selected such that an average distance between the representative speech segment and other speech segments within the same group is minimized.
7. The method of claim 1, wherein a number of the groups is determined based on at least one from a set of: a total number of speech segments, distances between the speech segments, and a desired reduction in concatenation cost data.
8. The method of claim 1, wherein the representative speech segment for each group is selected based on one of a median concatenation cost and a mean concatenation cost of each group.
9. The method of claim 1, wherein the speech segments include one of: individual phones, diphones, half-phones, and syllables.
10. A text to speech (TTS) synthesis system for generating speech employing compressed concatenation cost data, the system comprising:
a speech segment data store;
an analysis engine; and
a speech synthesis engine configured to:
determine a feature vector for each speech segment that comprises concatenation cost values of each speech segment with other speech segments;
apply distance weighting to each speech segment pair based on their respective feature vectors;
cluster the speech segments into a predefined number of groups such that an average distance between speech segments within each group is minimized;
select a representative speech segment for each group such that an average distance between the representative speech segment and other speech segments within the same group is minimized;
generate a compressed concatenation cost matrix based on the representative speech segments; and
pre-save the compressed concatenation cost matrix for real time computations in synthesizing speech.
11. The TTS system of claim 10, wherein the distance weighting is applied such that a sensitivity to compression errors is reduced.
12. The TTS system of claim 10, wherein the representative speech segment for each group is further selected based on center re-estimation.
13. The TTS system of claim 12, wherein the center re-estimation includes estimating a concatenation cost value based on a portion of whole samples such that a computation cost is reduced when speech segment numbers are relatively large.
14. The TTS system of claim 10, wherein the speech segment data store is configured to receive speech segments from at least one of: a user input and a set of pre-recorded speech patterns.
15. The TTS system of claim 10, wherein the analysis engine is configured to:
perform at least one from a set of: text analysis, prosody analysis, and phonetic analysis; and
provide input to the speech synthesis engine for segment selection based on the performed analyses.
16. A computer-readable storage medium with instructions stored thereon for generating speech employing compressed concatenation cost data, the instructions comprising:
determining feature vectors for speech segments based on a matrix of concatenation costs constructed along a preceding speech segments axis and a following speech segments axis;
applying distance weighting to each speech segment pair based on their respective feature vectors;
clustering the speech segments into M preceding segment and N following segment groups such that an average distance between speech segments within each group is minimized;
selecting a representative speech segment for each group;
generating a compressed concatenation cost matrix such that a concatenation cost between two speech segments is approximated by a concatenation cost between representative segments of respective preceding speech segment and following speech segment groups; and
pre-saving the compressed concatenation cost matrix for real time computations in synthesizing speech.
17. The computer-readable medium of claim 16, wherein the distance weighting is applied employing distance function:
\u03a3m=1n{abs(cci,m\u2212ccj,m)*K0\u2212(cci,m+ccj,m)}2, where cci,j are concatenation costs between speech segments i and j, and K0 is a predefined constant.
18. The computer-readable medium of claim 16, wherein the representative speech segment for each group is selected based on one of: minimization of an average distance between the representative speech segment and other speech segments within the same group, median concatenation cost of the group, and a mean concatenation cost of the group.
19. The computer-readable medium of claim 16, wherein the instructions further comprise:
determining M and N based on at least one from a set of: the total number of speech segments, distances between the speech segments, and a desired reduction in concatenation cost data
20. The computer-readable medium of claim 16, wherein a size of pre-saved concatenation data is reduced by n2(M\xd7N), where n is the total number of the speech segments.
The claims below are in addition to those above.
All refrences to claim(s) which appear below refer to the numbering after this setence.
1. An information recording medium having an information track formed spirally or in coaxial circles comprising:
a recordable area for information having a groove in a first depth being prerecorded with a frequency signal and a land pre-pit address signal from an inner circumference of said information track;
a first read only area having a pit in a second depth prerecorded with a frequency signal to be recorded with a reproduction signal as a pit; and
a second read only area having a pit in a first depth prerecorded with a frequency signal and a land pre-pit address signal to be recorded with a reproduction signal as a pit,
wherein a tracking error signal at a time of tracking off in a boundary between said first read only area and said second read only area is defined as a ratio of maximum amplitude in both directions from a center of maximum amplitude of said tracking error signal at the time of tracking off in said recordable area.
2. A reproducing method of information recording medium having an information track formed spirally or in coaxial circles comprising:
a recordable area for information having a groove in a first depth being prerecorded with a frequency signal and a land pre-pit address signal from an inner circumference of said information track;
a first read only area having a pit in a second depth prerecorded with a frequency signal to be recorded with a reproduction signal as a pit; and
a second read only area having a pit in a first depth prerecorded with a frequency signal and a land pre-pit address signal to be recorded with a reproduction signal as a pit,
wherein a tracking error signal at a time of tracking off in a boundary between said first read only area and said second read only area is defined as a ratio of maximum amplitude in both directions from a center of maximum amplitude of said tracking error signal at the time of tracking off in said recordable area,
said reproducing method comprising the steps of:
reproducing a land pre-pit address signal of said information track; and
tracking said boundary continuously on the basis of amplitude of a tracking error signal in said first read only area and said second read only area in accordance with reproduced said land pre-pit address signal.