In another session, the installation still going on, and no other errors are reported. Installation is still going and files are being copied from node1 to node2.
This issue seems to be coming from VMW and indicating that this is simply some performance (latency) hiccups.
Oracle RAC not starting upon rebooting. The following show some symptoms, tests, and what to look for.
Node2 not getting any RAC services status back when first startup or reboot.
[root@rac02 bin]# ./crsctl stat res -t
CRS-4535: Cannot communicate with Cluster Ready Services
CRS-4000: Command Status failed, or completed with errors.
Upon starting it all up, it has couple of errors. The "CRS-1705" and the "ora.diskmon" . Those are indications that there are some issues with the ASM storage being provisioned in node2.
[root@rac01 bin]# ./crsctl start cluster -all
CRS-2672: Attempting to start 'ora.cssd' on 'rac01'
CRS-2672: Attempting to start 'ora.diskmon' on 'rac01'
CRS-2672: Attempting to start 'ora.cssd' on 'rac02'
CRS-2672: Attempting to start 'ora.diskmon' on 'rac02'
CRS-2676: Start of 'ora.diskmon' on 'rac01' succeeded
CRS-2676: Start of 'ora.diskmon' on 'rac02' succeeded
CRS-2676: Start of 'ora.cssd' on 'rac01' succeeded
CRS-2672: Attempting to start 'ora.cluster_interconnect.haip' on 'rac01'
CRS-2672: Attempting to start 'ora.ctssd' on 'rac01'
CRS-2676: Start of 'ora.ctssd' on 'rac01' succeeded
CRS-2676: Start of 'ora.cluster_interconnect.haip' on 'rac01' succeeded
CRS-2672: Attempting to start 'ora.asm' on 'rac01'
CRS-2676: Start of 'ora.asm' on 'rac01' succeeded
CRS-2672: Attempting to start 'ora.storage' on 'rac01'
CRS-2676: Start of 'ora.storage' on 'rac01' succeeded
CRS-2672: Attempting to start 'ora.crsd' on 'rac01'
CRS-2676: Start of 'ora.crsd' on 'rac01' succeeded
CRS-1705: Found 0 configured voting files but 1 voting files are required, terminating to ensure data integrity; details at (:CSSNM00065:) in /u01/app/oracle/diag/crs/rac02/crs/trace/ocssd.trc
CRS-2674: Start of 'ora.cssd' on 'rac02' failed
CRS-2679: Attempting to clean 'ora.cssd' on 'rac02'
CRS-2681: Clean of 'ora.cssd' on 'rac02' succeeded
CRS-2672: Attempting to start 'ora.cssd' on 'rac02'
CRS-2672: Attempting to start 'ora.diskmon' on 'rac02'
CRS-2676: Start of 'ora.diskmon' on 'rac02' succeeded
CRS-2676: Start of 'ora.cssd' on 'rac02' succeeded
CRS-2672: Attempting to start 'ora.cluster_interconnect.haip' on 'rac02'
CRS-2672: Attempting to start 'ora.ctssd' on 'rac02'
CRS-2676: Start of 'ora.ctssd' on 'rac02' succeeded
CRS-2676: Start of 'ora.cluster_interconnect.haip' on 'rac02' succeeded
CRS-2672: Attempting to start 'ora.asm' on 'rac02'
CRS-2676: Start of 'ora.asm' on 'rac02' succeeded
CRS-2672: Attempting to start 'ora.storage' on 'rac02'
CRS-2676: Start of 'ora.storage' on 'rac02' succeeded
CRS-2672: Attempting to start 'ora.crsd' on 'rac02'
CRS-2676: Start of 'ora.crsd' on 'rac02' succeeded
Performing all the RAC-related storage checks will all appear hung. Once performing the oracelasm listdisks, it initiated the disk on node2 the RAC-related disk checks will show the disks output.
In the oracleasm log "/var/log/oracleasm" was showing "Disk "DISK01" does not exist or is not instantiated" when node2 rebooted. oracleasm configure showing disk scan is enabled. So, Scanning the disks manually after reboot seems to fix the issue. So, the issue has to do storage and timing. After some googling, my issue seems to match the following 2 notes from Oracle Metalink.
Oracle Linux 7: ASM Disks Created on FCOE Target Disks are Not Visible After System Reboot (Doc ID 2065945.1)
/usr/sbin/oracleasm.init" prior to scandisk and it solved my issue. That gives about 20 seconds for the storage to be presented before the scandisk and initiation.
With the 20 seconds delay, /var/log/oracleasm is showing "Instantiating disk "DISK01" upon reboot. The scandisk initialization attempt passes the Oracle RAC cluster start-up. If the disk is not initiated, nothing start-up in the cluster.
This blog is to explain I am getting the error when trying to create Oracle 19C database through DBCA and the mistake I made. The error is pretty straightforward and self-explanatory at the same time, it was vague.
When I saw the error, I thought, my grid_env or bashc profile had GI_HOME defined within. That was what was stated in Oracle Note (DBCA Fails to Start with DBT-10002 (Doc ID 2646840.1)).
As it turned out, I accidentally ran the DBCA on an absolute path in the Grid directory (/u01/app/version/grid/bin) instead of where my Oracle Software media directory. Once executing the DBCA from the media directory, it worked fine. Note: I have customized db_env, so, I use the absolute path for DBCA.
Upon rebooting after my newly set up Oracle RAC 19C, CRS isn't starting. After some basic checkings, I realized my "oracleasm listdisks" not show any disks at all on both nodes. It turned out the Openfiler has granted an accessible IPs have changed on my VMs. To correct the problem, I need to reset static IPs where Openfiler has granted access to where my ASM disk resides as SAN.
[oracle@rac02 ~]$ crsctl stat res -t
CRS-4535: Cannot communicate with Cluster Ready Services
CRS-4000: Command Status failed, or completed with errors.
[root@rac02 ~]# oracleasm listdisks
[root@rac02 ~]# oracleasm scandisks
Reloading disk partitions: done
Cleaning any stale ASM disks...
Scanning system for ASM disks...
[root@rac02 ~]# oracleasm listdisks
After setting the IPs to statis and restarted the network. Rescan the ASM disk on both nodes (yes, on both nodes. My second node not seeing the shared disk)
node1
[root@rac01 ~]# oracleasm listdisks
[root@rac01 ~]# oracleasm scandisks
Reloading disk partitions: done
Cleaning any stale ASM disks...
Scanning system for ASM disks...
Instantiating disk "DISK01"
[root@rac01 ~]# oracleasm listdisks
DISK01
[root@rac01 ~]#
node2
[root@rac02 ~]# oracleasm listdisks
[root@rac02 ~]# oracleasm scandisks
Reloading disk partitions: done
Cleaning any stale ASM disks...
Scanning system for ASM disks...
Instantiating disk "DISK01"
[root@rac02 ~]# oracleasm listdisks
DISK01
[root@rac02 ~]# oracleasm status
Checking if ASM is loaded: yes
Checking if /dev/oracleasm is mounted: yes
The second node needs to be manually started as well.
Use "crsctl start crs" or "crsctl start cluster -all"
or
[root@rac02 ~]# /u01/app/19*/grid/bin/crsctl start res ora.crsd -init
As the error described. It turned out my VIP IP is pingable which is not supposed to be. It should have been hidden. After double-checking a few rounds, I finally spotted that I made a typo of VIP IP and it is exactly the IP I had for the SCAN. Changing, it fixed my issue. So, look around, it may be somewhere in all the nodes /etc/hosts file that IP has been used and pingable.
This is a fairly generic share disk error. The reason I am getting this is when setting Oracle RAC on 2 VMs, I did not include ""diskLib.dataCacheMaxSize = "0" " to my VMX files on both VMs. So, upon starting up node2, it throws that error.
"The parent virtual disk has been modified since the child was created. The content ID of the parent virtual disk does not match the corresponding parent content ID in the child"
Resolution - add "diskLib.dataCacheMaxSize = "0" to the VMX file on all cluster nodes.
Adding a shared disk to the ASM disk group. Upon starting node2, it throws an error of "The process cannot access the file because another process has locked a portion of the file.". The reason for this is, VMX needs to enable disk sharing. This is only needed for adding a shared SCSI disk on VMware Workstation. Oracle Virtualbox has its own easier way to share disks.
Insert the following line to the VMX file.
"disk.locking = "FALSE"
As a side note for adding a share disk on Oracle Virtualbox, all needed are the following steps.
While installing the grid software, I encountered the issue of " INS-40912 Virtual host name xxxxx is assigned to another system on the network". It took me a while to realize, that I actually have the IP set on one of my eth interfaces as well as in the /etc/hosts file of the VIP IP section, it has been assigned to the VIP. So. the error did not lie about it. I had 2 IPs co-existed on the eth2 and also the /etc/hosts's VIP section. If you run into this issue, just make sure the IPs on the VIP section are not used or assigned to anything within the RAC or on the network. Normal ping on the IPs and name should help narrow down the issue. The chance of the later is low as during the setup, the installation will locate unused IPs for the VIP.
There might be a time that request to clone an extremely large database into an internal lab for data related troubleshooting. 99% of the time, development might only be interested in a certain set of tables but knowingly, a few tables are extremely large and unneeded. This can be easily achieved by performing a Datapump export from the customer database by excluding the table/s. In the following situation, Audit_Event data taking about 87% of the size of the database and they stored uninterested data.
expdp username/password dumpfile=mycloudexport.dmp directory=dumpdir_vcloud schemas=VCLOUD exclude=TABLE:\"IN \'AUDIT_EVENT\'\" logfile=vcloud_dumpfile.log By excluding the audit_event, the database pump process was speedy. Otherwise, the database backup file transfer will take over a day.
Default Bulk Insert do not work until I noticed the file wasn't what I saw from Notepad++. Upon, retrieving the file with Notepad, the data are all in 1 single line.
Upon firing up an old Oracle database that went down unexpectedly, I was getting the following error. It turns out the semaphore was too low. To temporarily changing it and starting it up.
SYS> startup ORA-27154: post/wait create failed ORA-27300: OS system dependent operation:semget failed with status: 28 ORA-27301: OS failure message: No space left on device ORA-27302: failure occurred at: sskgpsemsper
as root. echo "300 2000 200 128" > /proc/sys/kernel/sem
Trying to test new Oracle 12C feature "Moving or renaming datafiles" in attempt to fix "
ORA-19566: exceeded limit of 0 corrupt blocks for file". This Oracle error simply means, the block no longer belonging to any extents and Oracle 'marked' it as corruption, so, logically, if I can find a way to scrub it, it should fix the issue.
According to the Oracle DOC this should perform the following upon moving the datafile."When you rename or relocate online data files, the pointers to the data files, as recorded in the database control file, are changed. The files are also physically renamed or relocated at the operating system level." set current session to the database that I intend to test which is pdorcl.
SQL> alter session set container=PDBORCL;
Session altered.
Intent to change the vpxadmin3.dbf to vpxadmin4.dbf. Essentially, this can be point to somewhere else.
First time I get the chance to test out the "datafile move" feature. I can see that this may come very handy moving things around. Though, the performance can be questionable for large production environment.
This issue is caused by incorrect user being used during dbms_scheduler credential setup. The exit code of 7 is basically referring to permission issue. Pre-requisite should make sure the jssu, externaljob.ora and extjob have correct permission setup. Note: there are 2 ways to create scheduler credential. DBMS_SCHEDULER.CREATE_CREDENTIAL is deprecating in 12.1 and preferably be using DBMS_CREDENTIAL.CREATE_CREDENTIAL package 12.1 onward. Setup: Legit OS level user: oracle and password is vmware Instance level user setup: scott ascript.sh is a script that is calling another sql script at the OS level.
TEST # 1 : dbms_scheduler with valid OS user - PASSED
SQL> EXEC dbms_scheduler.run_job('SCRIPT_JOB3');BEGIN dbms_scheduler.run_job('SCRIPT_JOB3'); END;*ERROR at line 1:ORA-27369: job of type EXECUTABLE failed with exit code: 7 !@#--!@#7#@!--#@!&ORA-06512: at "SYS.DBMS_ISCHED", line 209ORA-06512: at "SYS.DBMS_SCHEDULER", line 594ORA-06512: at line 1
These are just some portions of my tests. There were other things I have tried such as incorrect passwords and etc and they were all failing with exit code of 7. At this point, I am not completely sure if this is a bug or part of the design.
RW-50010: Error: - script has returned an error: 1
RW-50004: Error code received when running external process. Check log file for details.
Running APPL_TOP Install Driver for dev instance
This error is very generic. One should not stopped at this level of troubleshooting. I do not believed looking at this one can guess what is the root cause. Quite a few Oracle noteid look the same but having a different root cause at the end of the resolution.
User will need to look at the install log to look at which stage and line that failed. At 29%, it does half a dozen things particularly moving java files around. User can look at the entire installation in the adrunfmw.cmd.
Locate the install log. It usually reside in the EBS install folder under temp. Example, oracle.apps.fnd.txk.install0.
Look for stdout and stderr or any exception errors. In my case, I encountered the following error.
<message>Process Completed (3) cmd /c rmdir /s /q C:\\oracle\\dev\\fs2\\FMW_Home\webtier\OPatch
Stdout: See C:\oracle\dev\fs2\EBSapps\appl\admin\dev_ebsdev\tmp\T1456139825284_122.tmp
Stderr: The system cannot find the path specified.
</message>
</record>
<record>
<date>2016-02-22T18:46:09</date>
<millis>1456141569147</millis>
<sequence>344</sequence>
<logger>oracle.apps.fnd.txk.install</logger>
<level>SEVERE</level>
<class>oracle.apps.fnd.txk.config.InstallService</class>
<method>printError</method>
<thread>10</thread>
<message>oracle.apps.fnd.txk.config.ProcessStateException: FileSys OS COMMAND Failed : Exit=3 See log for details. CMD= cmd /c rmdir /s /q C:\\oracle\\dev\\fs2\\FMW_Home\webtier\OPatch ## Node=NodeId=1698 Type=24 TypeName=filesys_patch_action Name= RefId=901 State=init ConfigDoc=APPS_OHS_HOME ParentDoc=null Topology=R12 Action=os_cmd
at oracle.apps.fnd.txk.config.FileSysPatchActionNode.doFileSysOSCmd(FileSysPatchActionNode.java:169)
at oracle.apps.fnd.txk.config.FileSysPatchActionNode.processState(FileSysPatchActionNode.java:101)
at oracle.apps.fnd.txk.config.PatchActionNode.processState(PatchActionNode.java:187)
at oracle.apps.fnd.txk.config.PatchNode.processState(PatchNode.java:338)
at oracle.apps.fnd.txk.config.PatchesNode.processState(PatchesNode.java:79)
at oracle.apps.fnd.txk.config.InstallNode.processState(InstallNode.java:68)
at oracle.apps.fnd.txk.config.TXKTopology.traverse(TXKTopology.java:594)
at oracle.apps.fnd.txk.config.InstallService.doInvoke(InstallService.java:224)
at oracle.apps.fnd.txk.config.InstallService.invoke(InstallService.java:237)
at oracle.apps.fnd.txk.config.InstallService.main(InstallService.java:291)
It is pretty self-explanatory. The required tmp file was not found. What I found though, not only the T1456139825284_122.tmp not found but the entire directory of C:\oracle\dev\fs2\EBSapps\appl\admin\dev_ebsdev\tmp\ was empty. The team have decided to get through this and re-visitng this missing file once the installation is successfully. Just to fast forward a little bit, there will be another missing file that I will point out in a bit.
Taking a look at the adrunfw.cmd command, I realized the script is hack-able. Since we are going to come back and revisit this issue if the installation ever completed, I decided to make a copy of the script and delete the script ending line so it break out of the loop if it failed and move on. Now - another issue found here, the j11067592_fnd.zip was not residing in the right path where the script will be copying from later on. It was sitting one level outside of fnd directory. I decided to just copy and paste the file into the fnd folder.
As expected, a few hours into the installation, it did make it passed 29% and completed successfully. As of the "T1456139825284_122.tmp" and possibly other files, I have no idea how to address it. That file was no where to be found in any of the medias nor I know what they are or for. Perhaps, run a one-off patching might restore them since, the it failed at the One-Off stage.
Please leave comments if you know what this T1456139825284_122.tmp is.