ccextractor/api/api_support.py

import ccextractor as cc
import ccx_to_python_g608 as g608
import python_srt_generator as srt_generator
##
#help_string to be included in the documentation
#it is the help string for telling what kind of output the user wants to STDOUT 
#the redirection of output to STDOUT can be changed easily to whatsoever usage the user wants to put in.
##
help_string = """
    Case is the value that would give the desired output.
    case = 0 --> print start_time,end_time,text,color,font
    case = 1 --> print start_time,end_time,text
    case = 2 --> print start_time,end_time,color
    case = 3 --> print start_time,end_time,font
    case = 4 --> print start_time,end_time,text,color
    case = 5 --> print start_time,end_time,text,font
    case = 6 --> print start_time,end_time,color,font
    """
text,font,color = [],[],[]
filename = " "
srt_counter = " "
def generate_output_srt(line):
    global text,font,color
    global filename, srt_counter
    if "filename:" in line:
        filename = str(str(line.split(":")[1]).split("\n")[0])
        #check for an alternative to wipe the output file in python
        fh = srt_generator.generate_file_handle(filename,'w')
        fh.write("")
        srt_generator.delete_file_handle(fh)
    elif "srt_counter-" in line:
        srt_counter = str(line.split("-")[1])
        fh = srt_generator.generate_file_handle(filename,'a')
        fh.write(srt_counter)
        srt_generator.delete_file_handle(fh)
    elif "start_time" in line:
        fh = srt_generator.generate_file_handle(filename,'a')
        srt_generator.generate_output_srt_time(fh, line)
        srt_generator.delete_file_handle(fh)
    elif "***END OF FRAME***" in line:
        data = g608.return_g608_grid(0,text,color,font)
        fh = srt_generator.generate_file_handle(filename,'a')
        srt_generator.generate_output_srt(fh, data)
        srt_generator.delete_file_handle(fh)
        text,font,color = [],[],[]
    else:
        g608.g608_grid_former(line,text,color,font)
Python bindings with extraction of CE608 grid and writing to a SRT output. (#768) * added python_extract to encoders_srt and the captions are being extracted in needed format. Search for an alternative to asprintf * Checking if the alternative to asprintf generate proper srts * CC captions accessible via python script * Removing python caption code from __wrap_write function * removing old cc_to_python functions * Removing python_subs structure and all the changes done for that struct * Removing filename functions from ccextractor.* * Renaming make_message to time_wrapper * Applying to python_extract codebase: SSA format * Added python_extract_time_based and done validation for ssa * pplying python_extract_time_based: Done validation for srt and webvtt * led attempt for SAMI support of python_extract. Code is commented * Appluing python_extract_time_based: validate support for SMPTETT * Added python_extract_transcript and made changes for time printing. * added show_extracted_captions_wtih_timings function * Added show_extracted_captions_with_timings to python script for testing purpose. * refactored extractors to api directory. commented out show captions in main() * build and build library working for the extractors. * made caption generator work with a 0.1 time sleep. Start refactoring * added asprintf for windows. * file being written in the running directory * Auto -deletion of python temporary file * Python captions printing status set to proper. * termination of tail successful * Writing successful for the sample * Generating unalternating output * adding api_support.py * Adding bld_flags in build_api * Added to build_library * Auto deletion of temporary file on SIGINT * Discussing Seg fault with Izaron * working for python and linux with samples. testing -out=pythonapi with stream * Done adding bitmap support * added -out=pythonapi support for bitmap * Setting the messages_target to 0 for output = pythonapi * Added wrapper for setting -out=pythonapi. Checking if -stdout value can be used in python. * adding the cc_to_stdout=1 value for -out=pythonapi. Thus generation of output file has been avoided. May be needed to change in future. * added extractor for g608 grid. removed sami extractor. need to work on overlap of -out=pythonapi and -out=g608 * Removed overlap of -out=pythonapi by adding -pythonapi and signal_python_api global variable. * added support for seperate c608 grid catching. Need to test the output via python. * added support for seperate printing of text font and color in CE608. Need to make sure that the function is inbuilt. * ADDED ce608 GRID SUPPORT FROM PYTHON need to discuss whether to keep the print_cc_grid function specific to the module or make it user accessible. Mostly it would be better to make it user accessible. * made changes in the call_from_python_api function such that only api_options is needed to be passed. An if statement before the call to g608_extractor has also been added. Waiting for Carlos to comment on the output generated till this stage. * added a signal_python_api check before calling every write function. Thus basic writing output can be avoided. * Commented all calls to python_extract_time_based. making changes to python_extract_g608 to be called only from the point when a g608 caption is detected. * Added pass_cc_buffer_to_python in encoders_common.c temporarily redefined get__encoded from static to normal included the above functions in encoders_common.h Added if-else statement for switch in encode_sub function. This is done mainly for making sure no output is generated in the api call. * Added ccx_encoders_python.c Defined pass_cc_buffer_to_python in ccx_encoders_python.c added if else statement in encode_sub's switch to make sure that the output is not generated in case of -pythonapi call * Removed __wrap_write from the entire code base. It's declaration and definition are only present in CCExtractor.* * Commented out the /dev/null part in ccx_encoders_common.c. Proceeding further on checking for file generation. * Added output_filename in array global variable and is generated in init_write function. included ccextractor.h in output.c to access global variable signal_python_api for avoiding output generation in init_write and invalid free in dinit_write. * Modified the definition of init_write function for accessing signal_python_api. * Deleted the commented part of /dev/null in ccx_encoders_common.c. * Added target_message=0 in -pythonapi param parsing in param.c to avoid the API from printing to STDOUT. Deleted the commented part of -out=pythonapi. Thinking of adding a different param for silencing the output when the call is made from python api. * Removed __wrap_write from ccextractor.c and ccextractor.h. * Added ccx_to_python_g608 and modified api_support.py file. added documentation in ccextractor.c. * added the generate srt script. However, some random characters are coming in first line. Need to talk about this. * Added SRT generator for python. Using string to remove the garbage value. Add code for srt counter and also the start_time and end_time conversion. * removed the trash characters and added code to print the timings. However, the last blank frame also results in a print. Need to take care of this. * rectified the mistake of writing only timings and not captions. now next step is to just make the timings print properly * some minor changes before diving into extracting srt_counter from the made codebase * Added extraction of srt_counter in python_extract via fflush srt_counter-value. Need to modify the processing in python. * Added the entire method to extract captions and generate srt files. Next, step would be a to define a concise function for writing the srt * Processing into a srt working properly. Next step is to add the information of font into the caption text. * the data is getting generated for proper SRT counters. * A turning point to the appraoch. Added END OF FRAME line for printing the data for every particular srt_counter. Proceeding further with the generation of srt by data manipulation. * some minor bugs but the output srt is being generated correctly. However, The font and colour encoding needs to be done. * Taken care of random characters. Need to discuss this with Carlos. Moving further to font/color processing. * Taken care of random characters. Need to discuss this with Carlos. Moving further to font/color processing. * Added fflush and cleaned up the python code of srt generation * Added <i> tag for italics. Proceeding further with other types. * Added the code to check for underline. However, need to check how CCExtractor generates srt when both italics and underline are present. For now a new line is added if both are present. * Shifting for making changes in th i/O work. * Stable ouput for samples with italics is being generated. * Added the PYTHONAPI macro definition and testing for its existence in the set_python_api function. * build script for linux is working correctly. Build_library is showing error of invalid def of set_pythonapi. Moreover, extractor has some memory seg fault. * Added mod to set a MACRO as my_python_api to set the callback function. Till now all calls to the reporter are commented. Working on getting the reporter to print the lines. * Changes have been implemented to bring reporter in working state. For now a constant string is passed from extractor. Need to make the proper parsing possible. * Changed the code in extractor such that entire grid is returned to the callback function. Need to provide this grid to the write function and also cleanup the codebase. * Writing the outputted srt in a file called "temp.srt". Need to modify init_write to push filename that is to be created in python using callback. * Added code to get start and end time simultaneously. entire SRT is getting generated. * removed ccx_python_encoders.c * Compiling and executing on Windows * Moved definitions get_line_encoded, get_color_encoded, get_font_encoded from ccx_encoders_g608.c to ccx_encoders_common.c. Also, deleted the static definition of get_font_encoded from ccx_encoders_webvtt.c * added a write statement in write_cc_bitmap_as_srt * Rectified transfer of get_line_encoded, get_color_encoded and get_font_encoded from ccx_decoders_common.c to ccx_encoders_common.c. 2017-08-20 15:54:35 +00:00			`import ccextractor as cc`
			`import ccx_to_python_g608 as g608`
			`import python_srt_generator as srt_generator`
			`##`
			`#help_string to be included in the documentation`
			`#it is the help string for telling what kind of output the user wants to STDOUT`
			`#the redirection of output to STDOUT can be changed easily to whatsoever usage the user wants to put in.`
			`##`
			`help_string = """`
			`Case is the value that would give the desired output.`
			`case = 0 --> print start_time,end_time,text,color,font`
			`case = 1 --> print start_time,end_time,text`
			`case = 2 --> print start_time,end_time,color`
			`case = 3 --> print start_time,end_time,font`
			`case = 4 --> print start_time,end_time,text,color`
			`case = 5 --> print start_time,end_time,text,font`
			`case = 6 --> print start_time,end_time,color,font`
			`"""`
			`text,font,color = [],[],[]`
			`filename = " "`
			`srt_counter = " "`
			`def generate_output_srt(line):`
			`global text,font,color`
			`global filename, srt_counter`
			`if "filename:" in line:`
			`filename = str(str(line.split(":")[1]).split("\n")[0])`
			`#check for an alternative to wipe the output file in python`
			`fh = srt_generator.generate_file_handle(filename,'w')`
			`fh.write("")`
			`srt_generator.delete_file_handle(fh)`
			`elif "srt_counter-" in line:`
			`srt_counter = str(line.split("-")[1])`
			`fh = srt_generator.generate_file_handle(filename,'a')`
			`fh.write(srt_counter)`
			`srt_generator.delete_file_handle(fh)`
			`elif "start_time" in line:`
			`fh = srt_generator.generate_file_handle(filename,'a')`
			`srt_generator.generate_output_srt_time(fh, line)`
			`srt_generator.delete_file_handle(fh)`
			`elif "*END OF FRAME*" in line:`
			`data = g608.return_g608_grid(0,text,color,font)`
			`fh = srt_generator.generate_file_handle(filename,'a')`
			`srt_generator.generate_output_srt(fh, data)`
			`srt_generator.delete_file_handle(fh)`
			`text,font,color = [],[],[]`
			`else:`
			`g608.g608_grid_former(line,text,color,font)`